Three Percent: Bloomberg Intelligence Says DeepSeek Has Closed the US–China AI Gap to a Record Low
DeepSeek's V4.1 Flash scored 81.1 on LiveBench — sixth globally, a record for a Chinese model — cutting the US performance lead to ~3% from 15% in January, and raising hard questions about export controls and the profitability of China's 1,100-model price war.
The single most consequential number in artificial intelligence this week is not a parameter count or a funding round. It is three percent. In a report published October 5, Bloomberg Intelligence (BI) senior analyst Robert Lea calculates that the gap between the best Chinese AI model and the best American one has collapsed to roughly 3% on LiveBench — the widest-read independent benchmark of model capability — down from about 9% in May and 15% at the start of the year. The proximate cause is DeepSeek’s V4.1 Flash, released in September, which scored 81.1 on LiveBench and ranked sixth globally: the highest position a Chinese model has held since the startup upended the industry with its R1 reasoning model in early 2025.
For two years, the operating assumption in Washington and Silicon Valley alike was that US export controls could hold Chinese frontier AI at arm’s length. A 3% delta is not arm’s length. It is rounding error on a leaderboard.
What the report actually says
Lea’s analysis, published October 5 and syndicated by the Straits Times, the Business Times, and Yahoo Finance, rests on LiveBench’s September standings. The scoreboard: Anthropic’s leading model holds the top score at 83.4, and DeepSeek V4.1 Flash sits at 81.1 — a 2.3-point spread that BI rounds to “roughly 3%.” That makes V4.1 Flash’s performance “comparable” to leading systems from Anthropic and OpenAI in Lea’s assessment, and marks the narrowest US–China gap BI has recorded.
The trajectory matters as much as the snapshot. The gap was about 15% earlier this year, 9% in May, and 3% now — a steady, monotonic squeeze rather than a one-off spike. And it has narrowed even as the absolute frontier moved: this is not Chinese models catching up to a static target, but closing distance on a target that is itself sprinting.
Two structural drivers sit behind the numbers, per the report. First is deepening AI expertise — China’s labs have spent two post-R1 years amassing frontier-scale research talent. Second, and more subversively, Chinese researchers have become expert at optimizing models for domestic hardware. When Nvidia’s best chips are restricted, you engineer your way around the restriction: sparse activation, aggressive KV-cache compression, quantization schemes tuned to the accelerators you can actually buy. DeepSeek’s V4 architecture paper reads as a manifesto for exactly this discipline.
BI is careful about what the number does not mean. Just three of LiveBench’s top 15 models are Chinese. Rankings shift month to month. And capability, as Lea notes, does not equal commercial gravity.
The model that moved the needle
DeepSeek V4.1 Flash is an awkward trophy for the “China is catching up” narrative precisely because it is not a maximalist flagship. It shipped September 10 with open weights under an MIT license, native vision, and a one-million-token context window. Its backbone is a 552-billion-parameter mixture-of-experts, but it activates only about 8B parameters while processing input and 16B while generating output — an asymmetric activation scheme that keeps inference radically cheap. The V4.1 update added a Causal Encoder-Decoder architecture and cut KV-cache memory to roughly a quarter of its predecessor’s requirement (around 890 bytes per token), which is what makes million-token context affordable to serve at all.
That engineering profile is the real story inside the story. A model that scores within 3% of Anthropic’s best while activating 1.5% of its parameters, under an MIT license anyone can download, is a direct attack on the economics of closed frontier APIs. It is the same playbook that made R1 a shock event in January 2025 — capability per dollar, not capability at any cost.
Export controls under strain
The geopolitical read is blunt. US chip restrictions were designed to curtail Chinese AI progress and deny Huawei and its peers the compute to build credible alternatives. Lea’s conclusion — that China’s progress “casts further doubt on the long-term sustainability of US technological supremacy in AI” — lands as export-control skepticism: if the gap can shrink from 15% to 3% in nine months under sanction, the controls are buying less strategic time than intended.
The report also notes the counterships. Chinese models face mounting US regulatory scrutiny and potential bans over model-distillation allegations — the accusation that Chinese labs train on outputs from US frontier models. And none of China’s labs have a clear path to profit: BI expects the Chinese AI industry to remain unprofitable until roughly 2030, trapped in low-margin token supply and a price war across a domestic market flooded with more than 1,100 large language models. ByteDance’s Doubao leads on app monetization; DeepSeek’s and Tencent’s chatbots remain free. In Lea’s framing, “a sustainable profit footing will require a cooling of competitive pressures, an industry shake-out and a more rational approach to pricing.”
Meanwhile the American labs are marching toward trillion-dollar valuations in anticipated stock debuts, with OpenAI and Anthropic explicitly marketing their capability edge. A 3% gap is still a gap — but it is a strange foundation for a supremacy narrative.
Why it matters
Three takeaways for anyone building on or investing in AI:
- Benchmarks are converging faster than business models. Chinese open-weight models now sit within noise of the closed frontier on composite scores. If capability parity holds, competition shifts to reliability, tooling, enterprise trust, and regulatory cover — domains where US labs still lead.
- Open weights are the asymmetric weapon. V4.1 Flash’s MIT license means the 81.1 LiveBench score is downloadable. Every enterprise weighing “closed best” against “open near-best, self-hosted, a fraction of the cost” now has a live option at 3% discount to the frontier.
- Policy has a math problem. If restrictions cannot preserve a meaningful capability buffer, Washington’s toolset narrows to adoption bans and alliance pressure — blunter instruments with collateral costs for US chipmakers, who lose the Chinese market either way.
The 3% number will be cited all week. The more durable signal is the slope: 15, 9, 3. Unless US frontier labs pull away again — and the last nine months suggest they have not — the “gap” framing itself may be approaching its expiration date.
Sources
- [1] https://www.bloomberg.com/news/articles/2026-10-04/us-lead-in-ai-over-china-narrows-after-deepseek-gains-bi-says
- [2] https://www.straitstimes.com/world/united-states/us-lead-in-ai-over-china-narrows-after-deepseek-gains-bloomberg-intelligence-says
- [3] https://finance.yahoo.com/technology/ai/articles/us-lead-ai-over-china-210300658.html
- [4] https://www.livemint.com/ai/deepseek-narrows-ai-gap-with-us-rivals-to-just-3-threatening-american-dominance-11791173222487.html
- [5] https://www.siliconflow.com/blog/deepseek-v4-1-flash-api-migration
- [6] https://www.mindstudio.ai/blog/deepseek-v4-1-flash-specs-architecture