Mercury 2.5 Preview Hits 1,107 Tokens Per Second: Inception's Diffusion LLM Quietly Rewrites the Economics of Fast Reasoning
Inception Labs' Mercury 2.5 Preview generates and refines tokens in parallel instead of one at a time, hitting 1,107 tokens/sec on standard GPUs with frontier-lite quality at a fraction of the price.
While the industry’s attention was fixed on Apple’s CEO transition and the G20 summit, Inception Labs shipped one of the most consequential model releases of the week. Mercury 2.5 Preview, which went live on OpenRouter on August 31, 2026, is the latest diffusion large language model (dLLM) from the Seattle-based startup — and it pushes the throughput frontier to a number that would have sounded like a typo a year ago: 1,107 tokens per second on standard commercial GPUs.
What makes a diffusion LLM different
Almost every language model you interact with today is autoregressive: it generates text one token at a time, left to right, each token conditioned on everything before it. The approach works remarkably well, but it imposes a hard ceiling on speed. If a response is 1,000 tokens long, the model must complete 1,000 sequential computational steps before you see the finished text. No amount of GPU parallelism inside a single forward pass can eliminate that dependency chain.
Diffusion LLMs break the sequence. Instead of writing tokens one at a time, Mercury generates an entire block of text in parallel and then iteratively refines it — the same fundamental idea behind image generators like Stable Diffusion, applied to language. Multiple denoising passes turn an initially rough block of tokens into coherent, high-quality output. The technique was long considered a research curiosity; Inception, founded by University of Washington professor Stefano Ermon and colleagues, turned it into a commercial product line.
The original Mercury paper, published on arXiv, demonstrated that dLLMs could sustain 1,100+ tokens per second on H100 GPUs in latency-optimized regimes while maintaining quality comparable to speed-optimized autoregressive models. Mercury 2, launched in February 2026, brought that architecture to enterprise workloads. Mercury 2.5 is the next step — and the first to put genuine reasoning capability in the sub-second latency band.
Mercury 2.5 by the numbers
The headline specs, drawn from Inception’s model page and OpenRouter’s provider listing:
- Throughput: 1,107 tokens/sec on standard NVIDIA GPUs, via parallel token generation rather than sequential decoding
- Context window: 260K tokens, up from 128K on Mercury 2
- Pricing: $0.20 per 1M input tokens and $0.75 per 1M output tokens, with cached input at just $0.02 per 1M
- Launch discount: 80% off via Inception on OpenRouter through September 8, 2026 at 07:00 UTC, bringing effective prices to roughly $0.04 in / $0.15 out
- Measured P50 on OpenRouter: 294 tok/s throughput, 1.14 s latency, with 99.99% provider uptime over three days
On raw economics, that discount pricing makes Mercury 2.5 Preview one of the cheapest reasoning-capable models on the market — undercutting models in its quality tier by a wide margin. Even at list price, the cached-input rate of $0.02 per 1M tokens is aggressive for agent workloads that repeatedly resend large system prompts and conversation histories.
The quality claim: 10+ points over Mercury 2
Speed without intelligence is a gimmick, and Inception knows it. The company claims Mercury 2.5 delivers a 10+ point jump in intelligence over Mercury 2, putting it in the same quality band as cost-optimized frontier models: OpenAI’s GPT-5.6 Luna (Low), Google’s Gemini 3.5 Flash-Lite, and Anthropic’s Claude Haiku 4.5. Those are precisely the models that power most high-volume production traffic today — the workhorses of customer support, enterprise search, and agent sub-tasks.
Feature-wise, Mercury 2.5 supports tunable reasoning levels (so developers can trade thinking depth for latency per request), parallel tool calls, and schema-aligned structured JSON output. Inception explicitly positions it for “production workloads where latency compounds: search agents, voice pipelines, and coding subagents.”
That last phrase deserves emphasis. In agentic systems, one user-visible task may fan out into dozens of model calls — planning, retrieval, tool use, verification. If each call takes 15 seconds, the total experience collapses. At 1,000+ tokens per second with sub-300ms time-to-first-token, diffusion models change the arithmetic of what an agent loop can afford. Voice is even more demanding: a model that “thinks” for ten seconds before speaking is unusable on a phone call, which is why Inception titled its Mercury 2 announcement “the first reasoning model fast enough to pick up the phone.”
Why throughput leaderboards don’t tell the whole story
A note of caution: benchmark followers should read the 1,107 tok/s figure correctly. It reflects the model’s parallel decoding capability in throughput-optimized conditions, while OpenRouter’s independently measured P50 sits at 294 tok/s — still exceptionally fast, and enough to place it at the top of the latency rankings, but an order-of-magnitude gap between marketing and measured reality. BenchLM, which tracks model evidence, currently lists Mercury 2.5 Preview with unconfirmed benchmark status, meaning third-party evaluations on standard suites are still catching up to the release.
The honest framing is this: Mercury 2.5 Preview is a preview, and its intelligence claims rest primarily on Inception’s own comparisons. But the speed is real, independently observable in production traffic, and the pricing is verified on OpenRouter.
The bigger picture: speed as a product category
Mercury 2.5 arrives at a moment when the industry is quietly bifurcating. Frontier labs compete on intelligence — bigger models, longer reasoning chains, higher benchmark ceilings. Meanwhile, an entirely different competition is happening below the frontier: who can deliver “good enough” intelligence at the lowest latency and cost. That second market is arguably larger. Most production tokens don’t need frontier reasoning; they need to be fast, cheap, and reliable enough to sit inside a product loop.
Inception’s bet is that diffusion is the architecturally superior answer to that market — and with Mercury now available on both its own API and OpenRouter, plus earlier availability on Azure AI Foundry, the company has built genuine distribution for a startup its size. Early OpenRouter traffic data shows real adoption, with top consuming applications already pushing over 200M tokens each.
Whether dLLMs eventually challenge the autoregressive orthodoxy at the frontier remains an open research question. But for the workloads that dominate real deployments — search, voice, subagents, autocomplete — Mercury 2.5 Preview has just moved the goalposts on what “fast” means. For the next week, at 80% off, it’s also one of the best bargains in AI.