From 6% to 67% in Seven Months: Chinese Models Now Route the Majority of Tokens on OpenRouter and Vercel
CNBC-reported platform data shows Chinese AI models handling 57-67% of OpenRouter tokens and 55% of Vercel traffic by August, as two House committees open investigations into the shift.
In February 2026, Chinese AI models processed somewhere between 6% and 13% of the tokens flowing through OpenRouter, the developer gateway that lets applications route requests across dozens of model providers. During the week of September 14, that share landed between 57% and 67%. On Vercel’s AI Gateway, a second major routing platform, Chinese models went from 11% of token volume in January to 55% by August. For the first time, Chinese-origin models now handle the outright majority of routed work on two of the developer ecosystem’s most-watched platforms — a shift reported by CNBC on September 26 and confirmed by figures both companies shared directly.
What the numbers actually measure
Both figures count tokens — the chunks of text that models read and generate — on platforms where developers deliberately choose among competing providers. They measure where work gets routed, not Chinese models’ share of all AI usage worldwide. That distinction matters: OpenRouter and Vercel are precisely the places where price-sensitive, high-volume engineering workloads congregate, so they capture the bleeding edge of the cost-driven substitution effect, not the average enterprise deployment.
Even so, the trajectory is stark. OpenRouter’s data covers companies in the United States, Europe, and 82 countries it groups as the “Global South.” Among Global South companies, Chinese models handled a full 67% of tokens in recent weeks. U.S. companies still account for roughly half of all tokens on the platform, which means the majority shift is not an artifact of geography — American engineering teams are participating in it too, even as U.S. frontier models continue to capture the majority of dollar spending.
That last gap is the story’s most important nuance: Chinese models dominate token counts while American labs still dominate revenue. Tokens measure volume; spending measures value. The platforms are telling us that cheap, capable open-weight models have captured the high-volume commodity tier of AI work, while the expensive frontier tier — the complicated tasks, the flagship deployments — still routes to U.S. labs.
Why the shift happened now
OpenRouter’s head of insights, Peter Walker, told CNBC that Chinese open-source models released this year have become credible for advanced multi-step agentic work, especially coding, in a way they simply were not in late 2025. That assessment — and it is an assessment, not a task-by-task benchmark result — pins the change on two converging forces.
First, capability convergence at the commodity tier. Releases from DeepSeek, Moonshot AI’s Kimi line, Alibaba’s Qwen family, and MiniMax closed enough of the gap on agentic coding workflows that “good enough” became the operative standard for production traffic. Vercel’s Harpreet Arora described the buying logic plainly: once a lower-cost model clears the quality bar for a job, the price difference becomes compelling. Companies still escalate complicated work to leading U.S. models — but the routine, high-volume middle of the workload distribution now has a dramatically cheaper alternative.
Second, pricing that undercuts by an order of magnitude. Chinese providers have priced API access 60-90% below comparable American models throughout 2026, and open weights add a third option entirely: self-hosting, which lets organizations run the architectures locally without transmitting proprietary data to external clouds. Rest of World reporting found software engineering teams confronting escalating compute expenses deliberately redirected less demanding workflows to Chinese systems to control burn rates.
Washington takes notice — with two committees and a sanctions threat
The commercial shift has arrived at exactly the moment U.S. policy is positioned to react badly to it. Two U.S. House committees are now investigating the impact of rising Chinese-model adoption among developers, according to CNBC. The investigations sit atop existing export controls that already restrict Chinese AI companies’ access to advanced chips, and they fold in two newer concerns: remote access to restricted compute through overseas data centers, and distillation — the technique of training a new model to mimic an established one’s outputs.
Treasury Secretary Scott Bessent has separately warned of potential sanctions over model distillation. Meanwhile, a countervailing camp of Silicon Valley investors and founders argues that restricting access to open weights would cripple domestic competitiveness, since American developers have built heavily on the same open ecosystems.
Daniel Remler of the Center for a New American Security framed the geopolitical stakes for CNBC: countries that routinize Chinese models could drift toward China’s technology sphere and, eventually, its geopolitical orbit. That remains a forecast rather than a finding — but it is the framing driving the Hill’s interest.
The uncomfortable arithmetic for U.S. labs
Put the two data series side by side and you get a squeeze. U.S. frontier labs keep the premium spending but are losing volume share at a pace that turned a single-digit percentage in February into a supermajority by mid-September. Chinese labs keep the volume — and with it, the production feedback loops, the developer mindshare, and the Global South relationships that compound over time.
The open-weight dynamic makes this harder to counter with product improvements alone. Every routing decision is a fresh auction, and the price gap is structural rather than promotional: it reflects genuinely lower inference costs and a strategic choice to trade margin for adoption. When Alibaba, Moonshot, and DeepSeek ship frontier-adjacent weights that anyone can self-host, the marginal cost of switching approaches zero for any workload that doesn’t strictly require a frontier model.
The South China Morning Post’s read, cited in the SiliconReport coverage, is that open releases from Moonshot and Alibaba have “narrowed the capabilities divide” — a polite way of saying the differentiation American labs can charge a premium for now lives in a narrower band of tasks than it did a year ago.
What to watch next
Three tests will determine whether this is a plateau or a trend line still climbing. First, whether Chinese providers convert token share into revenue share — if spending follows volume, the commercial pressure on U.S. labs becomes existential rather than uncomfortable. Second, whether Congress treats adoption itself as a security concern and moves from investigation to restriction, which would force an ugly choice between developer economics and policy compliance. Third, whether U.S. labs respond with their own structural price cuts — as OpenAI has already done piecemeal with GPT-6 Sol and Luna — or retreat to defending the premium tier.
The platforms, for their part, are neutral infrastructure: they route to whoever wins the auction. The week of September 14, that was Chinese models, two times out of three. Seven months earlier it was barely one time in ten. Whatever happens in Washington, that rate of substitution has already reset expectations for what AI inference should cost.