OpenAI's Ultrafast Mode Runs GPT-5.6 Sol at 750 Tokens per Second on Cerebras
OpenAI's new Ultrafast API tier, powered by Cerebras wafer-scale chips, serves its flagship GPT-5.6 Sol model at up to 750 output tokens per second — up to 14× standard speed.
On August 13, 2026, OpenAI opened a limited API preview of Ultrafast, a new service tier that runs GPT-5.6 Sol — the company’s most capable model — on Cerebras wafer-scale hardware. The headline number: up to 750 output tokens per second, roughly 14× the throughput of Standard processing, with OpenAI claiming no compromise on the model’s benchmark intelligence.
For context, frontier-class models typically generate somewhere between 50 and 100 output tokens per second on standard GPU infrastructure. At 750 tok/s, a full technical report that once took minutes to stream finishes before a developer has time to context-switch. OpenAI frames the change as order-of-magnitude — the kind of shift that doesn’t just speed up existing workflows but changes what products can be designed around model latency in the first place.
What Ultrafast Actually Is
Ultrafast is a tier within the OpenAI API, not a new model. The intelligence is still GPT-5.6 Sol; the difference is the serving layer underneath. Instead of GPUs, requests route to Cerebras systems built around the Wafer-Scale Engine (WSE) — a chip that is effectively an entire 300mm silicon wafer turned into a single processor, packing around 4 trillion transistors, 900,000 AI-optimized cores, and crucially 44 GB of on-chip SRAM.
That last spec is the whole story. Fast inference on large models is fundamentally a data-movement problem. On GPUs, model weights must shuttle repeatedly between on-chip memory and off-chip HBM to generate each successive token, and memory bandwidth becomes the bottleneck. Cerebras eliminates that movement entirely: weights stay resident on the wafer, and tokens flow uninterrupted through model layers pipelined across wafers. Because the weights never leave the silicon, the company argues the speed advantage scales naturally with model size — a direct bet that the approach holds as frontier models keep growing.
The tier is available initially to a select group of preview customers, with access expanding as capacity grows. OpenAI says early access has gone to companies across coding, e-commerce, financial research, and interactive production applications — named early users include Jane Street, Podium, Basis, and Rogo. OpenAI is also dogfooding the tier internally for incident-response log analysis, where minutes genuinely matter.
The Benchmark Claims — and Who Ran Them
Cerebras published several head-to-head comparisons worth taking with appropriate salt, since the company is explicit that these are its own measurements:
- Humanity’s Last Exam (HLE): GPT-5.6 Sol on Ultrafast answered all 2,500 PhD-level questions in 11 hours 11 minutes. Claude Fable 5 needed 78 hours 27 minutes — more than three days of continuous compute — to arrive at comparable conclusions. That is nearly a 7× end-to-end speedup on identical-difficulty work. (Benchmarked July 10 for Sol Ultrafast and July 13–15 for Fable 5, both on xhigh reasoning via Codex and Claude Code respectively.)
- GDP-Val: On this benchmark of economically valuable knowledge-work tasks, Ultrafast delivered a 5.6× end-to-end speedup over Standard processing with no quality degradation, tested July 31, 2026.
- Cross-vendor speed: Against output speeds reported by Artificial Analysis, GPT-5.6 Sol on Ultrafast runs 11× faster than Claude Fable 5 and 5× faster than Opus 4.8 on Fast mode.
These are vendor-run evaluations, not independent results — the 750 tok/s and 14× figures are OpenAI’s and Cerebras’s own numbers. But even discounted, the HLE result in particular illustrates a qualitative shift: frontier reasoning work that previously took a weekend of wall-clock time now fits inside a single working day.
Why Speed Changes the Product Surface
The interesting question isn’t whether 750 tok/s is fast — it’s what becomes buildable. Several categories immediately benefit:
Agents on the critical path. Until now, agent loops that reason for 30–60 seconds per step couldn’t sit inside interactive experiences or time-bounded operations. At 14× throughput, a model can spend three seconds on thinking tokens before producing output and still feel instantaneous to a user — reasoning depth stops costing latency budget.
Live incident response. When a production system is down and burning SLA minutes, root-cause analysis over thousands of log lines is exactly the workload where an hour of model time is unacceptable. This is why OpenAI itself points to outage triage as a first-party use case.
Financial research in motion. Jane Street’s John Crepezzi put it directly: “The increase in speed brought by Cerebras is impressive. It enables different ways of using the models, and makes it practical for developers to work in a more focused and productive way alongside them.” Markets don’t wait for batch jobs.
Voice and real-time interfaces. Human conversation pace is roughly 150 words per minute. Any latency budget that assumed 60 tok/s output constrained how much reasoning a voice product could afford per turn. At 750 tok/s, a model can think substantially harder mid-conversation without the interaction feeling broken.
The $10B Backstory
Ultrafast didn’t appear from nowhere. OpenAI tapped Cerebras for $10 billion in low-latency compute earlier in 2026, and this launch puts that capacity behind the flagship model rather than a smaller or specialized one. For Cerebras — a company that has spent years arguing its wafer-scale architecture is the right shape for inference while the market’s center of gravity sat firmly with GPU suppliers — landing the serving layer for OpenAI’s most capable model is the strongest production reference account it could ask for.
It also marks a notable crack in the NVIDIA-centric serving stack. OpenAI remains one of NVIDIA’s largest customers, but Ultrafast demonstrates that at least part of its frontier inference is now served on non-GPU silicon. Whether that remains a niche low-latency tier or expands as Cerebras capacity grows is the open question — OpenAI deliberately tied broader availability to capacity rather than announcing a fixed date.
Caveats and Open Questions
Three things temper the announcement. First, capacity is the near-term constraint: this is a limited preview, and neither company committed to general-availability timing. Second, pricing is unannounced — OpenAI has not published Ultrafast rates, and wafer-scale serving is unlikely to be cheap; the economics for commodity workloads may still favor Standard tier. Third, the quality-preservation claim rests on vendor benchmarks; independent confirmation will come as preview customers publish their own numbers.
The deeper question is architectural: Cerebras’s advantage rests on keeping model weights resident in on-chip SRAM, which works beautifully as long as frontier models fit in 44 GB per wafer (with techniques like FP8 compression). If model sizes outpace SRAM density gains, the pipelined-wafer approach faces the same memory-wall physics as everyone else. Cerebras’s counterargument is that its architecture scales smoothly with model size — a claim Ultrafast makes testable for the first time at frontier scale.
What to Watch
- Expansion pace. How quickly OpenAI widens the preview is the best proxy for whether Cerebras can deliver capacity at scale.
- Independent benchmarks. Watch for Artificial Analysis or preview customers to verify the 750 tok/s and quality-preservation figures.
- Competitive response. Groq, SambaNova, and NVIDIA itself (with low-latency serving stacks) all have incentives to answer. Inference speed is becoming a competitive axis distinct from raw intelligence — and that is arguably the most important signal in this launch: the frontier is no longer only about how smart the model is, but how fast it can think.
Sources
- [1] https://openai.com/index/previewing-ultrafast/
- [2] https://www.cerebras.ai/blog/accelerating-gpt-5-6-sol-ultrafast-with-openai
- [3] https://www.unite.ai/cerebras-runs-openais-gpt-5-6-sol-at-750-tokens-per-second-in-new-ultrafast-tier/
- [4] https://investors.cerebras.ai/news-releases/news-release-details/cerebras-powers-ultrafast-mode-openais-gpt-56-sol
- [5] https://news.ycombinator.com/item?id=49289844