← All posts / Tools

OpenAI's Ultrafast Mode Puts GPT-5.6 Sol on Cerebras Silicon at 750 Tokens Per Second

OpenAI's new Ultrafast API tier runs GPT-5.6 Sol on Cerebras accelerators at up to 750 output tokens per second — 14x standard throughput — collapsing the speed-versus-intelligence tradeoff for time-critical AI workloads.

OpenAI's Ultrafast Mode Puts GPT-5.6 Sol on Cerebras Silicon at 750 Tokens Per Second

OpenAI’s Ultrafast Mode Puts GPT-5.6 Sol on Cerebras Silicon at 750 Tokens Per Second

For years, AI builders have faced a stubborn tradeoff: you could have a fast model or a smart model, but rarely both at once. Larger reasoning models produce higher-quality output, but they pay for it in latency — users either wait minutes for excellent results or accept weaker answers delivered quickly. On August 13, 2026, OpenAI took a major swing at that tradeoff with the preview launch of Ultrafast mode, a new API service tier that runs its flagship GPT-5.6 Sol model at up to 750 output tokens per second — roughly 14 times faster than the standard tier — with, the company says, no compromise in intelligence.

The secret behind the speed is not a new model at all. Ultrafast is powered by Cerebras, the Silicon Valley chipmaker famous for its wafer-scale accelerators, which signed a ten-billion-dollar partnership with OpenAI earlier this year. The Ultrafast preview marks the most visible consumer of that deal to date: GPT-5.6 Sol’s weights served from Cerebras hardware rather than conventional GPUs, rewriting the economics of frontier-model latency.

What 750 tokens per second actually means

Numbers like “750 output tokens per second” can feel abstract, so it helps to ground them. A typical paragraph of prose runs about 100 words, or roughly 150 tokens. On Ultrafast, that paragraph arrives in under a fifth of a second. An extended analytical report of several thousand tokens — the kind of output that takes a standard-tier reasoning model a minute or more to stream — completes in seconds.

Cerebras put the tier through head-to-head benchmarking to quantify the gap. Running Humanity’s Last Exam, a 2,500-question benchmark designed to be answerable only by PhD-level experts across chemistry, economics, and literature, GPT-5.6 Sol on Ultrafast finished all questions in 11 hours and 11 minutes. Claude Fable 5, by comparison, needed 78 hours and 27 minutes — more than three days of continuous compute — to arrive at the same conclusions, making Ultrafast nearly 7x faster end-to-end at comparable accuracy. On GDP-Val, a benchmark for economically valuable knowledge work, Cerebras measured a 5.6x end-to-end speedup with no quality degradation.

According to output speeds tracked by Artificial Analysis, GPT-5.6 Sol on Ultrafast runs 11x faster than Claude Fable 5 and 5x faster than Opus 4.8 on its Fast mode. That positions Ultrafast as, in Cerebras’ phrasing, the world’s fastest frontier model.

Why this changes product design

OpenAI’s framing for Ultrafast is “more useful work per second” — the idea that frontier intelligence can now sit on the critical path of workflows where every second counts. The preview announcement sketches several concrete scenarios:

  • Incident response. Engineers could have production logs, code changes, and post-incident reports analyzed while an outage is still happening, helping pinpoint root cause and prepare a fix in real time. OpenAI says it already uses the model internally for exactly this purpose.
  • Finance. The model can evaluate market signals and flag suspicious transactions while conditions are still shifting, rather than delivering analysis after the moment has passed.
  • Customer support and e-commerce. Complex, multi-step inquiries can be resolved in real time; product questions, inventory checks, and personalized recommendations can complete before a hesitant buyer abandons their cart.
  • Research. Experiments that previously ran overnight as batch jobs become interactive work sessions — test an idea, review results, adjust, and kick off another run without breaking flow.

The productivity argument extends to agentic workflows. OpenAI researcher Jeffrey Wang described the experience bluntly: whereas formerly he might wait a couple of minutes for a task to finish, “it now finishes for me before I even have the opportunity to context-switch.” At 750 tokens per second, the model effectively keeps up with how fast a human can read, think, and redirect — collapsing the dead time that has defined working with reasoning models since they emerged.

Speed as a pricing lever

Beyond raw performance, Ultrafast reveals OpenAI’s evolving commercial strategy: monetizing inference speed in tiers. The company already offers a “Fast Mode” through the API promising up to 2.5x speed for GPT-5.6 Sol at roughly double the price. Ultrafast adds a third, faster, and — while pricing has not been announced — almost certainly pricier tier on top.

The logic mirrors how cloud providers like AWS have long charged premiums for higher-performance service classes. If latency becomes a genuine bottleneck across industries — and incident response, trading, and support use cases suggest it will — tiered inference gives OpenAI a direct cut of the revenue gains that faster AI creates. Compute speed becomes not just a technical metric but a pricing axis, the same way clock speed, memory, and network throughput became line items in enterprise infrastructure budgets.

Availability and caveats

Ultrafast is launching as a limited preview, initially available only through the OpenAI API for GPT-5.6 Sol and restricted to a select group of customers. OpenAI plans to expand access gradually as Cerebras capacity grows; interested organizations can sign up for updates through a form on the announcement page. No pricing or general-availability date has been disclosed yet.

Two caveats are worth keeping in mind. First, all benchmark comparisons above come from Cerebras and OpenAI themselves — independent third-party verification of the 750 tok/s figure under production workloads is still pending. Second, a limited-access, unpriced preview is a statement of intent more than a shipping product; whether wafer-scale serving can scale economically to broad API traffic remains the central question for Cerebras’ business.

Even with those caveats, the launch signals a meaningful shift in the frontier-model landscape. For most of the past decade, the industry’s race has been about intelligence — bigger models, better benchmarks. Ultrafast argues the next axis of competition is time: who can deliver frontier-quality reasoning fast enough to act on, in the moment it matters. With its Cerebras partnership now producing shippable product, OpenAI has drawn first blood in that race — and put the rest of the industry on notice that latency is now a frontier too.