← All posts / Research

SemiAnalysis: Open Models Now Close the Frontier Gap in Half the Time Every AI Era

A new SemiAnalysis study finds open-weight models are matching closed frontier models twice as fast with each successive AI era — from 12 months in the scaling era to under 5 months today, while Fireworks alone now serves 40 trillion tokens a day.

SemiAnalysis: Open Models Now Close the Frontier Gap in Half the Time Every AI Era

The analysts at SemiAnalysis have published one of the most cited studies of the week: “Are Open Models Catching Up?” Written by Evan Cloutier, Max Kan, Jordan Nanos, and Dylan Patel and released on August 21, 2026, the analysis reconstructs three years of model history and arrives at a deceptively simple finding — with each successive era of LLMs, open-weight models have taken half as long to catch up to the first closed frontier model of that era.

The timing matters. The past two months have been a breakout period for open source AI. Where January 2025’s “DeepSeek moment” produced headlines but little real usage, models like GLM 5.3 and Kimi K3 are now “genuinely capable of many of the same coding and agentic tasks that rocketed Anthropic to $65B+ ARR,” the authors write. Fireworks alone is processing over 40 trillion tokens per day — roughly double the OpenAI API’s volume at the end of March. This is no longer a curiosity; it is a structural shift with real economic workloads running on open weights.

Three eras, one accelerating trend

The study’s core methodological move is refusing to compare across eras with a single benchmark set. “Every benchmark is a product of a particular era,” the authors argue — benchmarks are created to discriminate between models of their time, hill-climbed until saturated, and then abandoned. Instead, they split LLM history into three eras and scored each with a curated benchmark composite, normalizing the best result in each era to 100.

Era 1 — Early scaling (2022–2024). The reference points are GPT-3.5 Turbo and Llama-2-70B, measured on GSM8K, HumanEval, TriviaQA, and MMLU-Pro. GPT-3.5 Turbo scored a normalized composite of 75.7; Llama-2-70B managed just 39.9 — a 35.8-point chasm. It took until Llama-3.1-405B in July 2024, roughly 12 months, for open models to close the gap. DeepSeek V3 then matched the era’s final frontier model, GPT-4o, in December 2024 with composites of 94.1 versus 95.5, while Qwen2.5-72B got within striking distance at a sixth of the parameter count.

Era 2 — Reasoning (2024–2025). OpenAI’s o1-preview shipped on September 12, 2024 and reset the benchmark landscape — grade-school math gave way to AIME and Humanity’s Last Exam. But this time the gap started much smaller: just 12.1 points, thanks to DeepSeek R1. The R1-0528 checkpoint closed it in May 2025 with a score of 78, an 8.5-month catch-up window.

Era 3 — Agentic (2025–today). With Claude Code’s general release in May 2025, the benchmarks that matter moved into the terminal: Terminal-Bench 2.1, BrowseComp-Plus, 𝜏³-banking, and DeepSWE. Most experts mark Opus 4.5 as the era’s start due to its reliability. Yet despite the enormous economic value at stake — and OpenAI and Anthropic compressing their release cadence to one model every 51 days on average, versus 213 days in Era 1 and 120 in Era 2 — the gap closed faster than ever. Kimi K2.6 surpassed Opus 4.5 with a score of 56.3 in 4.8 months, and GLM-5.2 cleared GPT-5.2 with a score of 72.4 in 6 months.

Twelve months, 8.5 months, under five. “The trend of the closing time halving with each subsequent era is remarkably consistent,” the authors note.

How they measured

The team ran most benchmark scores themselves using Prime Intellect’s evaluation stack — the environments hub and the evals harness from Prime-RL — with the remainder drawn from Artificial Analysis and Datacurve’s DeepSWE leaderboard. Open models were served the way they would have been at release: era-appropriate vLLM versions, hardware in use at the time, and model-card sampling settings. Closed models were evaluated against pinned API versions. Florian Brand (@xeophon) of Prime Intellect helped select benchmarks and verify correctness.

The caveats the authors themselves flag

The study is unusually honest about its limits. Benchmarks are not a perfect proxy for real work: “Kimi K3 may score higher than Fable 5 on our curated composite, but we still prefer using Fable at SemiAnalysis for our day to day work,” the authors admit — partly a tribute to how well Anthropic has productized Claude through Claude Code and Claude Tag, and partly an acknowledgment that public benchmarks can be gamed by building RL environments that closely mimic benchmark tasks.

They also preempt the objection that the Era 3 catch-up time is flattered by Western labs spending longer on safety testing than Moonshot or Zhipu. Not so, they argue: GPT-4 finished training 218 days before release, and even assuming Mythos finished training in mid-February, the delay before the Fable release was only 114 days.

Why this matters now

The study lands amid intense pressure on frontier-lab economics. Anthropic’s bankers are pitching an IPO at a valuation near $2 trillion on the strength of a $65B+ annualized revenue run rate, while OpenAI cut GPT-5.6 Sol API prices by more than 20% this week — its second pricing move in under a month — as capable open-weight rivals squeeze the frontier tier from below. If open models keep matching closed ones at a fraction of the cost, the obvious question is whether the model layer commoditizes, an outcome the authors acknowledge “would obviously be disastrous for frontier lab margins.”

Yet the report is less bearish on closed models than the headline trend suggests. The full forward projection is behind the study’s paywall, but the thrust is clear: each new era begins with a frontier research breakthrough that resets the gap, and the moat has increasingly shifted from raw model capability to the product wrapped around it. The gap closes; then someone invents the next era.

For everyone consuming tokens rather than selling them, the authors’ conclusion is the one to keep: “It is an exciting time to be a token consumer.”