← All posts / Models

8B Active Parameters and a 890-Byte Cache: DeepSeek's V4.1-Flash Outscores Its Own Flagship — and Phases It Out

DeepSeek's new 552B-parameter MoE activates just 8B parameters per input token, compresses its KV cache to 890 bytes per token, beats V4-Pro on Terminal-Bench 2.1 and DeepSWE — and inherits the flagship's API traffic on September 14.

8B Active Parameters and a 890-Byte Cache: DeepSeek's V4.1-Flash Outscores Its Own Flagship — and Phases It Out

8B Active Parameters and a 890-Byte Cache: DeepSeek’s V4.1-Flash Outscores Its Own Flagship — and Phases It Out

On September 10, 2026, DeepSeek released DeepSeek-V4.1-Flash, and the announcement read less like a product launch than a controlled demolition of its own flagship. The new model — the smallest in the company’s new architecture family — beat DeepSeek-V4-Pro on performance, cost, speed and total runtime in tests “by multiple parties,” DeepSeek wrote. The response: V4-Pro is being phased out. Starting at 04:00 UTC on September 14, all deepseek-v4-pro API requests will be routed to V4.1-Flash at Flash rates, until a future V4.1-Pro arrives. (Following user demand, the V4-Pro endpoint itself has since been kept alive with unchanged billing, but the direction of travel is unmistakable.)

The release is the sharpest expression yet of DeepSeek’s core thesis: that inference economics, not raw parameter count, decides who wins the agentic era. And the numbers behind that thesis are worth a close look, because they describe an architecture engineered with almost obsessive specificity for one workload — long-running AI agents.

The Causal Encoder-Decoder: asymmetric by design

V4.1-Flash is a multimodal Mixture-of-Experts model with 552B backbone parameters, organized as a 40-layer Transformer in a Causal Encoder-Decoder (CED) split: 20 causal-encoder layers followed by 20 decoder layers. The trick is what the decoder does with its KV cache. Instead of each decoder layer deriving KV states from its own hidden states, the decoder’s global KV cache is projected from the final encoder hidden states.

The payoff is asymmetry exactly where agentic workloads are asymmetric. During prefill — chewing through a long prompt, a repository dump, a conversation history — the model activates only 8B parameters per token. During decode — generating the actual response — it activates 16B. Agent workloads are input-heavy: a coding agent might read a million tokens of context to write a few thousand tokens of a patch. A model that spends 8B parameters on the input side of that exchange is priced structurally, not marginally, below models that burn full active-parameter budgets on both sides.

For comparison, DeepSeek-V4-Flash activated 13B parameters per token from a 284B backbone; V4-Pro activated 49B from 1.6T. V4.1-Flash is simultaneously bigger in total capacity (552B vs 284B) and cheaper per input token (8B vs 13B active).

Compressing the KV cache to 890 bytes per token

The headline engineering achievement is KV cache compression. For long-context agents, the KV cache — the memory of everything the model has already read — is the dominant cost. It consumes HBM during a run and often spills to SSD between steps, and cache-hit input charges are frequently the largest single line item in an agent’s API bill.

V4.1-Flash attacks this with a stack of techniques that the technical report describes with unusual precision:

  • Compressed Sparse Attention 2 (CSA2) assigns each attention layer one of three static modes — Full, Reindex, or Reuse — sharing main KV and indexer keys across layers and reusing Top-K sparse-attention indices. In the decoder, a Hierarchical Sparse Indexer restricts later indexing layers to a candidate pool built by the first Full Mode layer, bounding indexer cost independently of context length.
  • FP4 main KV caching in E2M1 format with one E4M3 scale per 16 channels squeezes the stored cache to near half-precision-quarter size.
  • SWA Bounded Replay reconstructs missing sliding-window KV states by replaying only the most recent n_win tokens — eliminating the need to persist SWA KV to SSD at all.

The combined result: a global KV cache footprint of 890 bytes per token — roughly one quarter of DeepSeek-V4-Flash, and (per the model card’s Figure 1) roughly a 437-fold reduction against DeepSeek-V1. In absolute terms, the announcement promises a KV cache needing 1/4 the HBM and 1/8 the SSD storage of the previous generation.

Context itself is 1M tokens, with sparse attention trained at 64K sequence length and extended to 1M at the 34-trillion-token mark of a 45T-token multimodal pretraining corpus. The model is natively multimodal: a from-scratch DeepSeek-ViT vision encoder (with 2D-RoPE and 3×3 pixel-unshuffle downsampling) and a two-layer MLP projector feed visual embeddings into the language model from the very start of pretraining.

Other components round out the design: Single-Pass mHC (a revised residual-stream mixing scheme with an efficient Mega-mHC kernel), an Engram conditional-memory module of 196B parameters accessed sparsely via token-based lookup, and DSpark speculative decoding with semi-autoregressive draft generation and confidence-scheduled verification. The MoE routing uses 1 shared expert and 384 routed experts per layer, activating 6 routed experts per token.

The benchmarks: small model, frontier results

On base-model knowledge and reasoning, V4.1-Flash mostly trades blows with the 1.6T-parameter V4-Pro — winning MMLU-Pro (74.1 vs 73.5) and BigCodeBench (60.6 vs 59.2), conceding SimpleQA-Verified and SuperGPQA. But the instruct-model results are where the story gets uncomfortable for the competition. At maximum reasoning effort, DeepSeek reports:

  • Terminal-Bench 2.1: 90.6 — above Opus-5.0 (89.1), GPT-5.6 Sol (88.8) and Kimi K3 (88.3)
  • DeepSWE v1.1: 74.2 — tied with Opus-5.0 (74.0), above GPT-5.6 Sol (73.0)
  • CyberGym: 88.1 — ahead of GPT-5.6 Sol (84.5) and GLM-5.3 (84.5)
  • Codeforces rating: 3471 — well above V4-Pro’s 3348
  • Agent’s Last Exam: 31.8 and AutomationBench: 54.8 — both category-leading in the table
  • HLE with tools: 63.9 — above Opus-5.0’s 63.6

The gaps run the other way on some harder agentic suites — Terminal-Bench 3.0 (30.0 vs Opus-5.0’s 43.3), Terminal-Bench 4.0 (31.2 vs 51.8), ProgramBench, ExploitGym — which keeps the frontier picture honest: this is a Flash-tier model winning on price-performance, not a clean sweep of the absolute state of the art.

The model also ships a continuously controllable reasoning effort from 1 to 100 — an integer dial trading inference cost for accuracy, a more granular alternative to the discrete thinking modes other labs offer.

Pricing: 50x cheaper at cache hit

The new pricing took effect at 04:00 UTC on September 10. For the deepseek-flash endpoint:

  • Input, cache hit: $0.006 per 1M tokens peak, $0.003 off-peak
  • Input, cache miss: $0.30 per 1M peak, $0.15 off-peak
  • Output: $1.20 per 1M peak, $0.60 off-peak

Peak windows are 01:00–04:00 and 06:00–10:00 UTC on weekdays; everything else is off-peak at half price. Context runs to 1M tokens with maximum output of 384K, and the concurrency limit is 2,500 — numbers that matter for the fleet-of-agents use case this model is plainly built for.

Set against Western frontier pricing — GPT-5.5 at roughly $5 per 1M input at the time of the V4 launch, Opus-class models far above that — the cache-hit rate of $0.006 is close to two orders of magnitude cheaper for the input-heavy traffic that dominates agent bills. Given that cache-hit charges are exactly what the 890-byte-per-token architecture attacks, the pricing and the engineering are two views of the same bet.

Open weights, and a deployment offer

True to form, DeepSeek open-sourced the model under an MIT license on Hugging Face and ModelScope, with 75,000+ downloads in the first days. The release notably skips a Jinja chat template, shipping instead a self-contained Python reference implementation plus deepseek-recipe — Rust libraries with Python bindings covering prompt encoding, tool calls, thinking mode, and streaming across Messages, Chat Completions and Responses APIs.

There is also a line in the announcement you don’t often see: “Planning a large-scale deployment with 2,000 GPUs + a storage cluster? Let’s talk.” DeepSeek is explicitly courting large self-hosters, and the emphasis on KV-cache SSD footprints suggests the company expects the model to be run in configurations where storage, not compute, is the binding constraint.

What it means

Three signals are worth taking away.

First, the era of “flagship” as a stable category is ending at DeepSeek. When a 552B model with 8B active prefill parameters beats the 1.6T flagship on the benchmarks that matter to agent builders, and the company responds by routing flagship traffic to the small model, the meaning of “Pro” has shifted from “most capable” to “most expensive per unit of capability.” DeepSeek has effectively repriced the frontier from above.

Second, the agent-cost war is being fought in the KV cache. ChatGPT-era economics were about cost per token generated. Agent-era economics are about cost per token re-read — and V4.1-Flash’s 4× HBM reduction, 8× SSD reduction and near-free cache hits are aimed squarely at that line item. Competitors pricing cache hits at hundreds of times DeepSeek’s rate will feel this in enterprise RFPs.

Third, open weights are now a delivery mechanism for architectural advantage, not just goodwill. The MIT license guarantees that vLLM, SGLang and the rest of the inference ecosystem will race to support CSA2 and CED — and every optimization they land makes DeepSeek’s architecture cheaper to run everywhere, including on clusters DeepSeek doesn’t own.

The caveats are real: self-reported benchmarks, a debut too fresh for independent verification, and clear losses on the hardest Terminal-Bench tiers. But the pattern by now is familiar. Every few months, the cheapest capable model in the world gets dramatically cheaper — and this time, it came with its predecessor’s retirement papers attached.