← All posts / Models

890 Bytes Per Token: How DeepSeek-V4.1-Flash Crushed KV Cache Costs and Retired Its Own Flagship

DeepSeek's new 552B-parameter open-weight model uses a Causal Encoder-Decoder architecture to cut KV cache to 890 bytes per token — 1/4 of the previous generation — while beating V4-Pro on agentic benchmarks, prompting DeepSeek to phase out its own flagship.

890 Bytes Per Token: How DeepSeek-V4.1-Flash Crushed KV Cache Costs and Retired Its Own Flagship

The most interesting model release of September isn’t the biggest one — it’s the cheapest one to serve. On September 10, 2026, DeepSeek released DeepSeek-V4.1-Flash, a 552B-parameter multimodal Mixture-of-Experts model with MIT-licensed open weights, a one-million-token context window, and a technical report whose subtitle tells you exactly what the company is optimizing for: Pushing the Limits of KV Cache Compression.

Four days later, DeepSeek did something almost unheard of in an industry that usually monetizes its flagships for years: starting at 04:00 UTC on September 14, every API request to deepseek-v4-pro — the company’s 1.6T-parameter flagship — began routing to V4.1-Flash instead, at V4.1-Flash prices. DeepSeek is phasing out its own flagship in favor of a model with roughly a third of the backbone parameters, because the smaller model wins on performance, cost, speed, and total runtime.

The problem: cache eats agentic workloads alive

To understand why this release matters, you need to understand where the money goes in modern AI inference. For chat, the expensive part is generating tokens. But for agentic workloads — coding agents, browsing agents, tool-using loops — the expensive part is reading. An agent re-feeds its entire conversation history on every step, which means the same prompt prefix hits the serving stack hundreds of times. Providers discount these cache hits heavily, but even discounted, cache-hit charges routinely account for the largest share of an agent’s total bill.

The size of that bill is driven by the KV cache: the per-token memory footprint the model maintains so attention layers can look back at earlier tokens. The bigger the KV cache, the more HBM you burn per concurrent user, the fewer users fit per GPU, and the more SSD storage you need to persist prefixes for reuse.

V4.1-Flash attacks this number directly. The model’s global KV cache footprint is 890 bytes per token — roughly one quarter of DeepSeek-V4-Flash (3,514 bytes) and, per the technical report’s own comparison, about a 437-fold reduction relative to DeepSeek-V1. At 890 bytes per token, a full one-million-token context needs only about 0.9 GB of KV cache. Persistent cache storage kept in SSD or host memory for prefix reuse falls to roughly one-eighth of the previous generation’s footprint.

How: a Causal Encoder-Decoder with three attention modes

The headline architectural change is what DeepSeek calls the Causal Encoder-Decoder (CED): a 40-layer Transformer organized as a 20-layer causal encoder followed by a 20-layer decoder. The trick is that the decoder’s global KV cache is projected from the final encoder hidden states rather than derived from each decoder layer’s own hidden states. The result is a genuinely asymmetric model: only 8B parameters activate per token during prefill (the input-heavy phase), while 16B activate during decode (the output phase). Input-heavy agentic workloads — precisely the ones drowning in cache costs — get the cheap path.

On top of CED sit two further compression mechanisms:

  • Compressed Sparse Attention 2 (CSA2) assigns each attention layer one of three static modes — Full, Reindex, or Reuse — sharing main KV and indexer K across layers and reusing Top-K sparse-attention indices. A Hierarchical Sparse Indexer in the decoder restricts later indexing layers to a candidate pool built by the first Full-mode layer, which bounds indexer cost independently of context length.
  • FP4 main KV caching quantizes the cache to E2M1 format with one E4M3 scale per 16 channels.
  • SWA Bounded Replay reconstructs missing sliding-window-attention KV states by replaying only the most recent n_win tokens, so SWA KV never needs to be persisted to SSD at all.

The rest of the stack reads like a compendium of DeepSeek’s research obsessions: a 196B-parameter Engram conditional memory accessed sparsely via token-based lookup (bringing total parameters to roughly 763B), Single-Pass mHC residual-stream mixing with a custom Mega-mHC kernel, and DSpark speculative decoding with confidence-scheduled verification. The MoE layers use 1 shared expert and 384 routed experts, activating 6 per token.

Vision is native rather than bolted on: a from-scratch DeepSeek-ViT encoder with 2D-RoPE and 3×3 pixel-unshuffle downsampling feeds a two-layer MLP projector, and visual embeddings were processed jointly with text from the very start of pre-training — a 45-trillion-token multimodal corpus, with sparse attention trained at 64K sequence length and context extended to 1M tokens at the 34T-token mark.

The numbers that justified killing a flagship

DeepSeek’s own benchmark tables show why the company was comfortable retiring V4-Pro. On base-model evaluations, V4.1-Flash-Base beats V4-Pro-Base (1.6T parameters, 49B active) on MMLU-Pro (74.1 vs 73.5), BigCodeBench (60.6 vs 59.2), HumanEval (79.4 vs 76.8), and GSM8K (93.0 vs 92.6) — while falling short on knowledge-heavy tests like SimpleQA-Verified (42.3 vs 55.2) and SuperGPQA (53.1 vs 53.9), a predictable trade for a third of the parameters.

The agentic results are where it gets striking. At maximum reasoning effort, V4.1-Flash posts 90.6 on Terminal-Bench 2.1 — ahead of Claude Opus 5.0 (89.1) and GPT-5.6 Sol (88.8) — plus 74.2 on DeepSWE v1.1 (vs 74.0 for Opus 5.0), 88.1 on CyberGym (best in table, ahead of GPT-5.6 Sol’s 84.5), 54.8 on AutomationBench, and 31.8 on Agent’s Last Exam, topping every listed frontier competitor on the latter two. It also leads Codeforces with a 3471 rating. The model still trails Opus 5.0 decisively on the hardest long-horizon terminal tasks (Terminal-Bench 3.0: 30.0 vs 43.3; Terminal-Bench 4.0: 31.2 vs 51.8) — the frontier hasn’t moved, but the mid-frontier just got dramatically cheaper.

One genuinely novel serving feature: reasoning effort is continuously controllable from 1 to 100 (an integer, not a low/medium/high toggle), letting developers trade inference cost against accuracy per request.

Pricing that reads like a provocation

The API pricing makes the cost story concrete. V4.1-Flash lists at $0.30 per million input tokens peak via DeepSeek’s own API, with off-peak rates at 50% of peak — and cached input at roughly $0.006–0.007 per million depending on provider. Third-party providers undercut even that: Fireworks lists $0.22/M input, $0.007/M cached input, $0.66/M output, with the full 1M context exposed. Context caching is where the architecture pays off — 2,500-concurrent-user capacity on the official platform is enabled by the 0.9 GB full-context cache footprint.

The model is live as deepseek-flash on the DeepSeek API, with native multimodal support. V4-Flash and V4-Flash-Vision-Exp are retired, their model names temporarily routing to V4.1-Flash for compatibility. Official partners WorkBuddy (including CodeBuddy) and OpenCode already support it. And in a line aimed squarely at hyperscale deployers, DeepSeek invites anyone “planning a large-scale deployment with 2,000 GPUs + a storage cluster” to get in touch — signaling the company expects the efficiency gains to translate into self-hosted deployments, not just API traffic.

The bigger picture

Two currents in late-2026 AI economics converge in this release. The first is that agentic workloads have inverted the inference cost model: input processing, not output generation, now dominates bills, and cache compression is therefore the highest-leverage optimization in the stack. The second is the open-weights pressure campaign: an MIT-licensed model that beats closed frontier models on Terminal-Bench 2.1, DeepSWE, CyberGym, AutomationBench, and Agent’s Last Exam — at $0.22–0.30 per million input tokens — is a direct attack on the pricing power of every closed lab.

There’s also a quieter signal for the industry: DeepSeek demonstrating that a well-executed smaller architecture can obsolesce a larger flagship within the same product family, five quarters after that flagship shipped. In a market where every lab is racing to scale, DeepSeek just made the counter-case — that the next competitive frontier is efficiency per byte of cache, not parameters per model. When V4.1-Pro eventually arrives, it will do so on top of this compressed-cache foundation, and the flagship tier will have to justify itself all over again.