The Sunset That Wasn't: DeepSeek V4.1-Flash Ships With 1M Context and $0.003 Cache Reads — and V4 Pro Gets a Reprieve
DeepSeek's new 552B-parameter open-weight model undercuts rivals on inference pricing just as the company walks back its plan to retire the V4 Pro API on September 14.
At 04:00 UTC on September 10, 2026, DeepSeek quietly rewrote the price floor of the frontier-model market. The company released DeepSeek-V4.1-Flash, a multimodal, open-weight Mixture-of-Experts model that posts benchmark numbers in the same neighborhood as considerably more expensive closed models — and it did so while attaching a price tag that reads like a typo. Then, two days later, it did something rarer still: it admitted a mistake. A scheduled shutdown of the popular V4 Pro API, due to take effect at 04:00 UTC on September 14, was cancelled after users pushed back.
What shipped on September 10
According to the official model card on Hugging Face, DeepSeek-V4.1-Flash is a multimodal Mixture-of-Experts model with 552 billion backbone parameters and support for contexts of up to one million tokens. The weights are available under the MIT license — among the most permissive licenses in existence — which means the model can be run, fine-tuned, distilled, and shipped inside commercial products with essentially no strings attached.
The architecture is where things get interesting. Third-party technical deep dives describe a causal encoder-decoder (CED) design with dual attention modes (dubbed CSA2) and modified residual connections, but the headline spec is the asymmetric activation: only about 8 billion parameters activate per input token, rising to roughly 16 billion during output generation. That split matters because it decouples the cost of reading context from the cost of writing answers. For agent pipelines that stuff enormous system prompts and retrieved documents into the context window but generate relatively short tool calls, the economics are dramatically better than a symmetric model of similar capability.
Independent trackers also report the model was trained on roughly 45 trillion tokens, and that it accepts images natively — vision input is built into the base model rather than bolted on through a separate encoder product. One wrinkle that sparked debate in the open-source community: an audit of the released safetensors files argues the true stored parameter count is closer to 750 billion, because much of the model is held in 4-bit precision and different counting conventions produce different totals. DeepSeek’s official documentation stands by the 552B backbone figure, but the discrepancy has become a lively thread on r/LocalLLaMA — a fitting illustration of how opaque “parameter count” remains as a comparison metric in 2026.
The pricing that made people look twice
New pricing took effect at 04:00 UTC on September 10, and it is structured around a simple idea: off-peak rates are exactly 50% of peak rates. The numbers, per million tokens:
- Input (cache miss): $0.30 peak / $0.15 off-peak
- Input (cache hit): $0.006 peak / $0.003 off-peak
- Output: $1.20 peak / $0.60 off-peak
That $0.003 cached-input rate is the number that made the rounds on developer social media, and VentureBeat’s coverage framed the release around it. Cached input — the repeated system prompts and static context that agent frameworks resend on every call — is the silent line item that bankrupts naive RAG deployments, and DeepSeek has now priced it at a level where re-sending a 100K-token prompt a thousand times a day costs about thirty cents. Early benchmark reports claim the model’s agentic scores surpass the V4-Pro-Preview tier and compete with closed flagships like GPT-5.6 Sol and Claude Opus 5 on coding and tool-use evaluations, at a small fraction of their per-token cost. Treat vendor-adjacent benchmark claims with the usual skepticism, but the independent leaderboard coverage has been broadly consistent with DeepSeek’s own numbers.
The retirement that wasn’t
The original plan was ruthless in a very Silicon Valley way. When V4.1-Flash launched, the older V4 Flash was retired immediately, and DeepSeek announced that after 12:00 Beijing time (04:00 UTC) on September 14, traffic sent to the higher-tier V4 Pro would be rerouted to V4.1-Flash. On paper this was reasonable — the new model beats the old one on most published benchmarks — but V4 Pro had become a production workhorse, and its behavior on edge cases is something enterprises had spent months calibrating around. A forced swap, executed on a four-day deadline, is the kind of change that breaks agent pipelines in subtle ways: slightly different tool-call formatting, slightly different refusal boundaries, slightly different math.
Users complained, loudly and specifically. And on September 12, DeepSeek’s API changelog was updated with a sentence that startled anyone used to the take-it-or-leave-it deprecation notices of American cloud providers: “In response to user demand, we have decided to continue providing API services for DeepSeek V4 Pro after September 14, 2026, with the billing method remaining” unchanged. The sunset became optional. As of this writing, the September 14 switchover deadline is hours away, and V4 Pro customers who changed nothing will simply keep working.
Why this matters beyond one model
Two things are worth taking away.
First, the pricing structure is a strategic weapon, not a promo. DeepSeek has consistently used off-peak discounting to smooth its own inference load, and the 50%-of-peak rule extends that discipline to customers. When a frontier-adjacent open-weight model charges $0.60 per million output tokens off-peak, every closed-model vendor’s price list becomes a negotiation waiting to happen. The efficiency race among Chinese labs — squeezing capability per FLOP rather than bragging about total compute — is no longer a curiosity; it is the reference curve the rest of the market gets measured against.
Second, the reprieve is a small but real data point about how the open-weights economy disciplines vendor behavior. DeepSeek cannot fully lock customers in, because the weights of its previous models are already in the wild; anyone truly desperate can self-host. That fact makes purely involuntary migrations harder to impose, and it means listening to users is not just good manners but good business. A deprecation notice from a closed-model provider is a fait accompli; from an open-weight lab, it is an opening bid.
For developers, the practical guidance is straightforward: V4.1-Flash is now the default recommendation for new builds — the combination of MIT weights, one-million-token context, native vision, and near-free cache reads is hard to beat for agent workloads. V4 Pro users got their stay of execution, but the direction of travel is obvious, and the smart move is to begin testing the migration now, on your own schedule, rather than waiting for the next sunset date to be announced.
The model weights are on Hugging Face, the pricing is live, and the deadline that dominated developer forums all week turned out to be bendable. In a market defined by escalating capability claims, the most subversive thing DeepSeek shipped this week may have been a changelog entry that simply said: we heard you.
Sources
- [1] https://www.deepseek.com/en/news/deepseek-v4-1-flash/
- [2] https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash
- [3] https://api-docs.deepseek.com/updates/
- [4] https://venturebeat.com/technology/deepseek-v4-1-flash-debuts-with-0-003-1m-off-peak-cached-input-rate-and-benchmarks-eclipsing-gpt-5-6-sol-claude-opus-5
- [5] https://apidog.com/blog/deepseek-v4-pro-retirement-migration-guide/
- [6] https://www.datacamp.com/blog/deepseek-v4-1-flash
- [7] https://llm-stats.com/models/deepseek-v4.1-flash