← All posts / Industry

The Great AI Price Collapse: US Labs Slash Token Prices 25% in a Month as Chinese Models Take Over

Prices for leading US AI models have fallen nearly 25% since mid-July as OpenAI and Anthropic fight a two-front war against cheap Chinese rivals and their own ballooning inference costs — and enterprise buyers are switching sides.

The Great AI Price Collapse: US Labs Slash Token Prices 25% in a Month as Chinese Models Take Over

For most of the past three years, the price of frontier AI moved in one direction: up. Each new generation of models came with bigger context windows, longer reasoning chains, and bills that made CFOs flinch. That era just ended — abruptly. According to the Financial Times, prices for leading US AI models have fallen nearly 25% since mid-July, the sharpest one-month decline since the launch of ChatGPT, and the cause is not generosity. It is competition from China.

What the numbers say

The FT’s analysis, published August 14, tracks list prices across the major US labs and finds a coordinated slide. The headline moves:

  • OpenAI cut GPT-5.6 Luna by 80% — from $1.00 to $0.20 per million input tokens and from $6.00 to $1.20 per million output tokens. GPT-5.6 Terra fell 20% in the same July 30 announcement. Frontier-tier Sol pricing held unchanged.
  • Anthropic launched Claude Opus 5 at roughly half the price of its frontier model Fable 5 — $5 per million input and $25 per million output tokens, versus Fable 5’s $10/$50.
  • The blended result: a basket of leading US model prices is down almost 25% in a single month, per the FT tally — and enterprise AI costs have hit their 2026 low, according to separate research covered by SCMP.

To put the new floor in perspective: GPT-5.6 Luna now outperforms Claude Fable 5 on several benchmarks at about one-sixteenth the cost, by OpenAI’s own accounting. Mid-tier intelligence that cost $50 per million output tokens a year ago now costs $1.20.

Why now: the China squeeze

The proximate cause is a demand-side shift that would have been unthinkable eighteen months ago. Data from OpenRouter — the API aggregator that routes a huge share of developer LLM traffic — shows Chinese models have overtaken Claude and ChatGPT in token volume among developers.

The milestones came fast this summer:

  • In July, Chinese models accounted for 57% of tokens used by US companies on OpenRouter during one week — the first time China’s share crossed the US share since the data began.
  • By mid-August, DeepSeek alone captured more than 27% of total processing volume on OpenRouter, overtaking Google’s ~25% share.
  • CGTN-reported OpenRouter rankings in early August showed Chinese models holding the top five spots in weekly token usage — 28.13 trillion tokens for Chinese models versus 4.38 trillion for US models, a streak then in its 14th week.

The economics driving the switch are stark. DeepSeek’s V3.2 charged $0.42 per million output tokens while Anthropic’s Claude Opus charged $75 — a 178x gap. Moonshot’s Kimi K3, one of the strongest Chinese agents models, lists at $15 per million output tokens, roughly half of OpenAI’s GPT-5.6 Sol. Even DeepSeek’s brand-new flagship V4-Pro, with its controversial peak-hour pricing, lands at $1.32/$3.96 per million tokens — far below Western frontier tiers.

Enterprise buyers are voting with tokens

The FT’s earlier reporting on the same trend named names: DoorDash, Siemens, and Airbnb are all running trials or production workloads on Chinese models as inference bills mount. Fortune’s July survey of the phenomenon found companies drawn not just by price but by a performance gap that has narrowed to near-parity on many enterprise tasks — summarization, classification, extraction, translation — even where Chinese models still trail on the hardest reasoning benchmarks.

This is the part that should worry San Francisco more than the benchmark charts. Enterprise AI adoption was supposed to be a moat: lock in developers with the best model, and they’ll pay frontier prices forever. Instead, OpenRouter’s routing data suggests developers treat models as interchangeable commodities below the frontier tier, switching to whatever delivers acceptable quality per dollar. When your product is a utility, price is the product.

The margin math nobody wants to say out loud

Price cuts of this magnitude collide head-on with the industry’s other existential problem: inference costs are exploding as agents run for hours and reasoning models “think” in thousands of tokens before answering. The same FT piece notes that labs are simultaneously trying to protect the revenue growth stories that underpin eye-watering valuations — OpenAI’s reported $40 billion annualized run rate, Anthropic’s $11.5 billion Q2 — while gross margins on inference compress with every price cut.

There are only three ways this resolves. Labs can drive inference costs down faster than prices fall (the Cerebras partnership behind OpenAI’s new 750-tokens-per-second Ultrafast mode is exactly this play). They can move value up the stack into agents, subscriptions, and advertising — OpenAI’s coming ads for European free-tier users and its enterprise revenue crossover both fit. Or they can keep cutting and absorb the margin hit, betting that scale and cheaper silicon eventually catch up. All three are happening at once, which is a sign of an industry that doesn’t yet know which lever works.

The open-weights undertow

One more structural force is pulling prices down: open weights. Alibaba’s Qwen family just passed 3 billion downloads in six months, dwarfing Google’s 418 million and Meta’s 227 million for all of 2026, and models like Qwen 3.8 27B now ship Apache 2.0 with frontier-adjacent coding scores. When a developer can host a 27B model that hits 61.7 on SWE-Bench Pro on their own GPU, the pricing power of closed APIs erodes from below — and Chinese labs, whose strategic bet on openness was once dismissed as a necessity, now set the global price floor.

US labs still hold the frontier. GPT-5.6 Sol, Claude Fable 5, and Gemini 3.7 Pro remain the models you reach for when the task is genuinely hard. But the frontier is a small and crowded market, and the volume — the hundreds of trillions of tokens that make up the real business of AI — is rapidly becoming a commodity market priced in Shenzhen and Hangzhou rather than San Francisco.

The 25% monthly price collapse isn’t a sale. It’s the sound of the AI industry’s business model adjusting to what it has actually become: infrastructure.