DeepSeek V4-Pro-0813 Goes GA: 1.6T Open-Weight Reasoning Flagship Challenges Closed Models
DeepSeek's 1.6T-parameter MoE flagship exits preview with a 96.4% SWE-bench score, MIT open weights, and pricing 57x cheaper than top closed models.
On August 12–13, 2026, DeepSeek moved its V4-Pro model from preview to general availability with the 0813 build, cementing its position as the most formidable open-weight AI lab in the world. The release brings a 1.6-trillion-parameter Mixture-of-Experts model to production APIs, benchmarks within fractions of a point of the best closed models on multiple leaderboards, and does it all at a fraction of the cost — all under an MIT license.
What Shipped
The DeepSeek V4-Pro-0813 build is the general availability release of the V4-Pro family. The headline specifications are straightforward but staggering:
- Total parameters: 1.6 trillion (1.6T)
- Active parameters per token: ~49 billion via MoE routing
- Context window: 1 million tokens
- Max output: up to 384,000 tokens
- Training data: over 32 trillion tokens, trained with FP4 precision
- License: MIT (open weights)
- API pricing: $0.435 per 1M input tokens, $0.87 per 1M output tokens
The model is accompanied by its smaller sibling, V4-Flash, which runs 284B total parameters with ~13B active per token — designed as a high-speed, low-cost tier for everyday workloads. Together, the two models form a routing strategy: push bulk traffic to Flash, escalate hard reasoning tasks to Pro.
The architecture uses a hybrid attention stack that DeepSeek says reduces FLOPs to roughly 27% and KV cache memory to about 10% compared to dense models of equivalent scale. This efficiency is what makes a 1.6T model commercially viable to serve at 83.2 output tokens per second with a 1.63-second time-to-first-token — speeds that rival much smaller models.
Benchmark Performance: Closing the Gap
The 0813 build’s most striking achievement is how close it lands to the very top of every major benchmark, despite being fully open-weight and radically cheaper.
On SWE-bench Verified — the gold standard for software engineering agents — V4-Pro-0813 scores 96.40%, placing second overall on the vals.ai leaderboard and within just 0.60 points of the closed-weight leader. That puts it ahead of Claude, GPT-5, and other major closed models on this particular evaluation.
On the Artificial Analysis Intelligence Index, the model scores 53, ranking #2 out of 104 tracked models. Its overall conversational benchmark score of 87.9 on community leaderboards places it within a tenth of a point of the top model (Fable 5 at 88.0) — while costing roughly 57x less on output pricing.
For coding-specific evaluations, V4-Pro-0813 scored 71.6% on Aider (a real-world code editing benchmark), slightly edging out competing models. The deepSWE benchmark clocked in at 62.7%. These are numbers that, just a year ago, were the exclusive domain of frontier closed models from OpenAI and Anthropic.
The Architecture Story
DeepSeek’s MoE approach is the core enabler. Rather than activating all 1.6 trillion parameters for every token, the model routes each token through only ~49 billion parameters — about 3% of the total. This is why a model with more total parameters than GPT-5 or Claude Opus can still serve at competitive speeds and prices.
The FP4 training precision represents another efficiency frontier. By training in 4-bit floating point rather than the more common BF16 or FP8, DeepSeek claims substantial reductions in training compute and memory bandwidth requirements. Combined with the hybrid attention mechanism — which sparsifies the attention computation to reduce both FLOPs and memory — the architecture is engineered for inference economics as much as for raw capability.
The 1M-token context window with up to 384K tokens of output is another differentiator. Most frontier models cap output at 16K–64K tokens. The 384K ceiling makes V4-Pro-0813 viable for long-form document generation, multi-file code refactoring, and extended agentic workflows where a single response might need to produce an entire codebase or report.
Open Weights: The Strategic Implications
The MIT license is perhaps the most consequential detail. Unlike Meta’s Llama series, which carries acceptable-use restrictions, or Google’s Gemma, which has commercial caveats, DeepSeek V4-Pro ships with a permissive license that permits unrestricted commercial use, modification, and redistribution of the full 1.6T model weights via Hugging Face.
This matters for several reasons. First, enterprises concerned about vendor lock-in can self-host a frontier-tier model — something that was economically impractical before MoE made 1.6T-parameter models servable on reasonable hardware. Second, researchers gain full access to inspect, fine-tune, and build upon a genuinely top-tier model rather than a distilled approximation. Third, the open-weight release creates a forcing function: competitors like Anthropic and OpenAI can no longer claim that open models are meaningfully behind closed ones.
The Competitive Landscape
DeepSeek positions V4 at roughly the level of Gemini 3.1, GPT-5.4, and Claude Opus 4.6 — the previous generation of frontier models. While newer closed models like GPT-5.6 and Fable 5 maintain small leads on certain benchmarks, the gap has compressed to single-digit percentage points in most categories, and V4-Pro wins outright on SWE-bench Verified.
The pricing differential is where DeepSeek’s advantage becomes acute. At $0.435/$0.87 per million tokens, V4-Pro costs roughly 3–7x less than comparable closed models on input and up to 57x less on output for certain workload tiers. For high-volume agentic applications — where models may generate millions of tokens per task — this order-of-magnitude cost reduction changes what is economically feasible.
What to Watch
The 0813 release is not without caveats. Independent benchmark verification is still underway — many of the headline scores come from vendor-reported or community-collected data rather than fully controlled third-party evaluations. The model is text-only, lacking the multimodal vision and audio capabilities that newer closed models offer. And DeepSeek’s position as a Chinese company means the open weights carry geopolitical considerations for some enterprise customers, particularly in U.S. defense and government contexts.
Nonetheless, the trajectory is clear. With each DeepSeek release, the practical gap between open and closed frontier models narrows further. The 0813 build is the strongest evidence yet that open-weight AI can match closed models on the benchmarks that matter — while radically undercutting them on price. For developers building AI-powered products, that is an equation worth paying attention to.
Sources
- [1] https://openrouter.ai/deepseek/deepseek-v4-pro-0813
- [2] https://glm5.app/blog/what-is-deepseek-v4-pro
- [3] https://artificialanalysis.ai/models/deepseek-v4-pro
- [4] https://vals.ai/benchmarks/swebench
- [5] https://www.techtimes.com/articles/324241/20260813/deepseek-v4-pro-0813-goes-ga-benchmark-claims-await-independent-proof.htm
- [6] https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro