The headline for DeepSeek-V4.1-Flash is a 552B-parameter multimodal MoE with a million-token context, but the number that matters for anyone running agents is 890 bytes of KV cache per token — about a quarter of the previous V4-Flash. When your workloads are long-horizon agents that reread a growing context every step, the KV cache, not the weights, is what fills your HBM and caps your batch size. The full tech report sits on the HF paper page.

The early open-source reaction tracks the same split I’d flag internally. Context Studios calls the compression groundbreaking and a top open-weights result, while MindStudio walks through the CSA2-plus-FP4 mechanics without the hype. KDnuggets plays the useful skeptic: it argues rivals like GLM-5.3-Flash post stronger overall numbers for less, so the win here is the serving architecture, not the leaderboard. That is the right frame — you reach for V4.1-Flash when your constraint is KV-cache memory under long-context agent load, not when you are chasing the top reasoning score.

tags: [ ai-infrastructure ] [ llm-ops ] [ agentic-ai ] [ research ]