The Flash That Ate the Flagship: DeepSeek V4.1 Flash Arrives — and Retires V4 Pro
DeepSeek's new 552B MoE with native multimodality beats its own flagship on agent benchmarks — so the company is routing all V4 Pro traffic to it, at Flash prices, from September 14.
Model releases usually follow a script: a bigger flagship arrives, and last year’s model gets a price cut and a polite demotion. DeepSeek just tore up that script. On September 10, 2026, the company officially released DeepSeek-V4.1-Flash — the smallest model in its new architecture family, with native multimodal visual understanding — and in the same breath announced that its own flagship, V4 Pro, will stop serving its own traffic. From 12:00 Beijing time (04:00 UTC) on September 14, every request to deepseek-v4-pro will be routed to V4.1 Flash and billed at Flash prices, until a future V4.1 Pro arrives.
For anyone building on the Pro endpoint, that is not a routine deprecation notice. It is a forced migration to a cheaper, faster model — one that DeepSeek claims, with its own benchmark tables, has comprehensively surpassed the flagship on performance, cost, speed, and total task-completion time.
What V4.1 Flash actually is
The headline numbers describe a 552-billion-parameter Mixture-of-Experts backbone — nearly double the 284B of the previous V4 Flash — but the active-parameter story is what makes it runnable: the model activates only 8B parameters to read input and 16B to generate output, down from 13B across the board on V4 Flash. The architecture has changed shape too. Instead of a pure MoE decoder, V4.1 Flash uses a Causal Encoder-Decoder design with 20 encoder and 20 decoder layers, paired with an “Engram” conditional-memory module of 196B parameters that sits separate from the backbone.
Two other facts stand out. Vision is no longer a bolted-on experimental sibling (the retired V4-Flash-Vision-Exp): V4.1 Flash is multimodal from the start of pre-training, with native image understanding. And the release ships under an MIT license, continuing DeepSeek’s habit of giving away frontier-adjacent weights.
The efficiency engineering is the quiet star of the release. Compressed Sparse Attention 2 and FP4 KV caching push the global KV cache down to 890 bytes per token — roughly a quarter of V4 Flash’s footprint and a claimed 437-fold reduction against DeepSeek-V1. In a 1M-token context window, that is the difference between a memory budget that fits on a rack and one that does not.
The benchmark picture — and its asterisks
DeepSeek’s changelog lists an unusually broad eval sweep: GPQA Diamond 90.9, HLE 36.8 (39.1 on the pure-text subset), Codeforces rating 3471, MathArena Apex 65.6, Terminal-Bench 2.1 at 90.6, DeepSWE v1.1 at 74.2, CyberGym 88.1, and HLE-with-tools at 63.9.
The comparison against the outgoing flagship is where it gets interesting. On agentic work, Flash wins outright: 90.6 vs 87.9 on Terminal-Bench 2.1, and 74.2 vs 62.7 on DeepSWE v1.1 — an 11.5-point gap on one of the hardest software-engineering agent benchmarks in circulation. On unaided knowledge and reasoning, the flagship still holds ground: V4 Pro’s GA numbers show 92.4 on GPQA Diamond and 55.2 on SimpleQA-Verified for the base models, against Flash’s 90.9 and 42.3. In other words, the new model is a better agent but a slightly worse encyclopedia.
DeepSeek also did something unusually honest for a model card: it published the same model scored through eight different agent scaffolds. The spread on DeepSWE from tooling alone was 8.7 points — wider than the gap between the top three models in any published comparison. Anyone who reads single-number benchmark tables as ground truth should sit with that figure for a moment.
Pricing, migration, and what it means
V4.1 Flash costs $0.30 per million input tokens and $1.20 per million output tokens at peak, with off-peak rates halved. For traffic migrated from V4 Pro, that works out to roughly a 77 percent cut in cache-miss input cost and 70 percent in output cost. On the API, the model is served under the name deepseek-flash; the retired deepseek-v4-flash and deepseek-v4-flash-vision-exp names are temporarily routed to the new model so existing code keeps running.
The release follows a characteristically compressed timeline. A limited-time beta appeared on September 9 under a model name that literally encoded its expiry — deepseek-v4.1-flash-expires-on-0910 — with beta pricing matching V4 Flash and a 20-concurrent-request cap. One day later, the formal launch landed, complete with the V4 Pro routing decision. DeepSeek’s Harness tooling was updated to 0.1.5 the same day, adding V4.1 Flash support, file uploads, and sidebar previews.
The bigger signal is architectural. DeepSeek is explicit that the new design targets “a higher capability ceiling, faster inference, higher throughput, and scaling to larger models” — which reads as a dress rehearsal for a V4.1 Pro built on the same Causal Encoder-Decoder plus Engram skeleton. When the smallest model in a new family already displaces the previous flagship on agent work at a fraction of the cost, the family’s ceiling moves somewhere interesting. The open-weight ecosystem gets MIT-licensed weights on Hugging Face, and every API customer gets a four-day countdown to decide whether their Pro-dependent pipeline is ready.
For a company that has spent two years resetting expectations about what frontier-adjacent capability should cost, retiring your own flagship by routing it to the cheap model is the most DeepSeek move possible.