← All posts / Industry

56% and Rising: Open-Weight Models Just Took Over the Majority of Production Tokens

Vercel's September AI Gateway Production Index shows open-weight models running 56% of all gateway tokens in August — the first majority in the index's history — while price per token fell 23.2% in a single month.

56% and Rising: Open-Weight Models Just Took Over the Majority of Production Tokens

For eight months, the shift had been building one percentage point at a time. In December 2025, open-weight models processed fewer than one in ten tokens crossing Vercel’s AI Gateway. In August 2026, they crossed the line: 56% of all gateway tokens ran on open-weight models — the first month they carried more production volume than every closed-weight lab combined.

The number comes from Vercel’s AI Gateway Production Index for September 2026, published September 17. The gateway routes tens of trillions of tokens per month between production applications and AI labs, which makes this one of the few datasets that reflects what teams actually deploy — not what they say they deploy in surveys.

The crossover, in context

The index traces an unbroken climb: open-weight share rose every single month from April through August, going from 13% to 56% of total volume. December-to-August, that is an eightfold gain in share. The trajectory matches what other observability points have hinted at all summer — Vercel’s own July index had open-weight models at 29% of tokens on under 4% of spend, with DeepSeek alone moving nearly a quarter of all volume.

The spend picture is the other half of the story. Open-weight models ran 56% of August’s tokens but accounted for only 14% of spend. The frontier labs still hold the majority of dollars — and within that premium tier, the money is consolidating rather than dispersing.

Deflation is now structural

The clearest consequence shows up in pricing. The average price per token across the gateway fell 23.2% in August, the third consecutive monthly drop and the steepest since April. Among teams running more than ten million tokens in both July and August, the median team paid 7.6% less per token — more than double July’s 2.9% decline.

Read that carefully: the median team got almost 8% more inference for the same budget in a single month. That is the mechanical effect of capable open-weight alternatives sitting one API call away. Teams are increasingly reserving frontier-priced models for the tasks that genuinely justify the premium and routing everything else down-market.

Fable’s fall and the loyalty question

The most striking case study in the report is Anthropic’s own model ladder. Fable 5 is Anthropic’s most capable and most expensive model; Opus 5 sits one tier below at roughly half the price per token. When Fable 5’s access was restored on July 1 after a US export control was lifted, it surged to 13.2% of gateway spend. Then Opus 5 came online at the end of July — and in August, Fable 5’s spend share collapsed to 4.9% while Opus 5’s rose to 22.5%. Nine in ten teams running Fable cut their usage, and more of them moved to Opus 5 than to any other model.

Anthropic, notably, kept 64% of all gateway spend in August, and has held at least 61 cents of every gateway dollar every month since December. Its models have occupied the top two spend spots every month — even as the models in those slots changed. Customers left the flagship, but they stepped down inside the same vendor.

The report frames this as the index’s central lesson: loyalty follows the model profile, not the lab. When a successor preserves what users valued — same workloads, better price — the lab keeps the customer. Z.ai’s GLM-5.3-Flash overtook GLM-5.2 one day after appearing on the gateway and was processing two-thirds of Z.ai’s tokens by August 31. When the successor doesn’t offer a relative edge, customers walk: Gemini 3 Flash has lost 95% of its token share since May, and more than three-quarters of that volume went to other labs. Google’s share of gateway token volume fell from 30% to 5%, with Gemini 3 Flash alone accounting for 22 of the 25 lost points. Roughly half the departed volume went to cheaper models led by GPT-5.6 Luna; the other half went up-market to Claude Opus 5 and Sonnet 5.

Astra’s opening salvo

The September index also carries an early read on the newest frontier contest. OpenAI’s GPT-6 Astra launched on the gateway September 3 at the same price as Anthropic’s Fable 5.1 (which had launched two days earlier). Within 48 hours, Astra accounted for one in every three dollars spent on OpenAI models through the gateway. Over each model’s first twelve days, Astra took 7.7% of all gateway spend — more than double Fable 5.1’s 3.7% — and was used by twice as many teams, passing Fable 5.1’s cumulative spend on day four.

OpenAI’s lineup is splitting into a volume tier and a premium tier: Astra and GPT-5.6 Sol processed 27% of OpenAI’s tokens but 71% of its spending from September 4–16, while Luna and Nano moved more than twice the tokens for about a ninth of the spend. Luna alone processed eight times Astra’s tokens; Astra still generated four times Luna’s revenue.

And adoption is still accelerating: per a September 18 addendum, the newly arrived Jev decision model became the fastest-adopted model in gateway history, reaching nearly 13% of paid teams within 24 hours.

Why it matters

A 56% open-weight majority does not mean the frontier is dying — it means the frontier is being redefined as a premium niche while a competent, commoditized middle absorbs the volume. For engineering teams, the deflation numbers are effectively a standing discount: routing tiers by task difficulty is now worth real money at any scale beyond hobby use. For labs, the Vercel data argues that pricing ladders and successor-model continuity matter as much as benchmark scores: Anthropic lost its flagship’s share and kept the revenue; Google kept iterating and lost the traffic. The next index, with a full month of Astra-versus-Fable 5.1 and post-crossover open-weight data, may show whether 56% was a peak or just a waypoint.