The headline for DeepSeek-V4.1-Flash is a 552B-parameter multimodal MoE with a million-token context, but the number that matters for anyone running agents is 890 bytes of KV cache per token ā about a quarter of the previous V4-Flash. When your workloads are long-horizon agents that reread a growing context every step, the KV cache, not the weights, is what fills your HBM and caps your batch size. The full tech report sits on the HF paper page.
- šÆ KV cache is the agentic bottleneck, not FLOPs ā long-horizon tool use makes workloads input-heavy, and prefill plus a fat KV cache is where the cost actually lives.
- ā” Two levers stacked: cross-layer KV reuse in Compressed Sparse Attention 2, plus FP4 KV caching, get the resident footprint down to 890 bytes per token.
- š Asymmetric compute: the Causal Encoder-Decoder activates 16B params per decode step but only 8B during prefill ā cheaper exactly where agent context ingestion hurts.
- š Off-accelerator footprint drops to roughly 1/8 of V4-Flash through SWA Bounded Replay, which matters once you page context to SSD or host memory.
- š” The production read: a smaller KV cache buys more concurrent agent sessions per GPU, or longer context at the same budget ā the lever that moves serving cost, not benchmark bragging.
The early open-source reaction tracks the same split Iād flag internally. Context Studios calls the compression groundbreaking and a top open-weights result, while MindStudio walks through the CSA2-plus-FP4 mechanics without the hype. KDnuggets plays the useful skeptic: it argues rivals like GLM-5.3-Flash post stronger overall numbers for less, so the win here is the serving architecture, not the leaderboard. That is the right frame ā you reach for V4.1-Flash when your constraint is KV-cache memory under long-context agent load, not when you are chasing the top reasoning score.
tags: [ ai-infrastructure ] [ llm-ops ] [ agentic-ai ] [ research ]