Squeezing Tokens: How Chinese AI Labs Are Closing the Gap by Being More Efficient
A new Fortune analysis argues China's labs are matching US frontier AI not by outspending it but by out-engineering it — cheaper attention algorithms, open weights, and enterprise adoption data that show the gap has narrowed to 2.7%.
For two years, the received wisdom about the US-China AI race was simple: America trains the frontier, China copies it. A widely shared analysis published by Fortune on September 13, 2026 pushes back on that framing with an uncomfortable argument — Chinese labs are not merely catching up by distilling American models. They are closing the gap because, starved of compute, they were forced to become better engineers of efficiency, and that discipline is now winning them enterprise customers in the United States itself.
The context: an intelligence race measured in percentages
How close is close? By one widely cited Stanford benchmark aggregation earlier this year, Anthropic’s top model led DeepSeek’s best by just 2.7%. That is the margin separating “months ahead” from “neck-and-neck,” and it frames every other number in this story. When the performance delta shrinks to single digits, the decision matrix for a CTO shifts from “which is better?” to “which is better per dollar?” — and that is precisely the terrain where Chinese labs have concentrated their advantage.
The immediate news hook is diplomatic. Earlier in the week, the FBI, NSA, and CISA jointly published advisory AA26-251A alleging that six Chinese AI companies — DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI — bought bulk subscriptions to American frontier APIs and extracted “capabilities worth billions” by training on the outputs since 2024. The advisory claims this method let DeepSeek understate its famous $5.6 million training cost. China’s foreign ministry called the accusations “groundless” and attributed the country’s progress to “high-level scientific and technological self-reliance.”
Both things can be true at once. Distillation may explain part of the parity. But analysts interviewed by Fortune argue the deeper advantage is architectural — a culture of doing more with less that US labs, floating on hyperscaler budgets, never had to develop.
Attention, engineered under constraint
The technical core of the argument is the attention mechanism — the 2017 Google innovation that lets a language model weigh every token against every other token to decide which relationships matter. Attention is what makes long context possible, and it is also what makes it brutally expensive: computational cost grows with the square of the context length.
Brendan Burke, semiconductors and supply-chain analyst at Futurum Group, told Fortune that Chinese labs “found algorithms that reduce the complexity of those calculations by an order of magnitude, and then achieve better results because they’re able to summarize the most relevant tokens.” The constraint was imposed from outside: US export controls cut China off from Nvidia’s best accelerators and pushed the ecosystem toward domestic alternatives like Huawei — hardware that, on paper, cannot match the clusters US labs command. The White House’s own numbers make the asymmetry stark: the US holds roughly 74% of world compute capacity, reinforced by hyperscaler data center buildouts measured in gigawatts.
“Because they had less compute to work with, they found that computationally efficient method instead of just throwing more compute at an inefficient technique, as U.S. labs initially did,” Burke said. US frontier models, by contrast, can afford to be “token hogs” — systems designed to reason exploratorily, burning tokens to test hypotheses, because the compute behind them is abundant.
Anyone following model releases this year has seen the artifacts of this discipline. DeepSeek’s V4 line introduced compressed-attention techniques claiming a 73% reduction in per-token inference FLOPs. Z.ai’s GLM-5.3-Flash cut attention compute by 3× against its own predecessor while shipping open weights under an MIT license. Alibaba’s Qwen3.8 family activates a sliver of its total parameters per token through aggressive mixture-of-experts routing. None of these are clones of American architectures; they are divergent engineering lineages shaped by scarcity.
The enterprise receipts
Efficiency claims are cheap. Adoption data is not. The Fortune piece assembles several:
- Larridin, an AI measurement platform, tracks enterprise engineering workflows and finds that Chinese models like GLM 5.2 and Kimi 2.6/2.7 handle roughly 75% of engineering tasks “reasonably well” at one-fifth the cost of US models. Frontier US models retain the edge on the most complex tasks — but the majority of everyday enterprise engineering no longer requires the frontier.
- Hugging Face reported that Chinese open-source models accounted for 41% of total downloads last year — a larger share than US models.
- Ramp’s AI index shows the share of businesses paying for platforms with access to open-source and Chinese-developed models rising from 4.5% in January to 6.1% by July.
- Named adopters now include DoorDash (CEO Andy Fang called Moonshot’s Kimi “cheaper” and “better quality”), Cursor (which used Kimi to help build its Composer 2 agent), Airbnb and Siemens (experimenting with Alibaba and DeepSeek models), and Thomson Reuters, which adapted Alibaba’s open-source Qwen into an in-house model — Thomson-1 — to handle document review previously done by Claude.
The open-weight distribution model deserves much of the credit. DeepSeek’s R1 and its successors can be downloaded from Hugging Face, fine-tuned privately, and served through US-based clouds like AWS — a path that neutralizes data-sovereignty objections and lets wary American enterprises adopt Chinese model technology without sending a byte to a Chinese company.
The counterweight
The story has a sober counterweight, and Fortune gives it space. US labs remain months ahead on the hardest tasks. Mike Finley, CTO of AnswerRocket, frames the frontier as the “existence proof” the rest of the industry builds against: “The work they do would simply not be possible without the frontier labs blazing the trail.”
There is also the nagging question of provenance. If the distillation allegations in AA26-251A hold up, some fraction of Chinese model capability is derivative of American training runs — meaning the efficiency story and the extraction story are entangled rather than competing explanations. Efficiency explains how Chinese labs serve capable models so cheaply; distillation allegations, if proven, would explain part of how the capability got there. Regulators will have to untangle the two before writing policy that assumes a clean separation.
What it means
The efficiency thesis matters because it changes what “winning the AI race” means. If capability parity is within a few percent and the cost gap is 5×, then the battleground shifts from raw intelligence to economics at scale — and the side optimized for scarcity is structurally positioned for a world where, per McKinsey’s survey, 20% of business leaders already say token costs are constraining their AI adoption.
For US labs, the uncomfortable implication is that abundance was never free. Compute wealth bought speed, but it also bought habits — exploratory reasoning, generous context, token-hungry agentic loops — that a price-sensitive market may not reimburse forever. For Chinese labs, the implication is the mirror image: the constraint that defined them is becoming their export product. The 2.7% question is no longer who is ahead. It is who can afford to stay in the race.