OpenAI Launches Ultrafast Mode: GPT-5.6 Sol at 750 Tokens Per Second via Cerebras
OpenAI is previewing Ultrafast, a new API service tier that runs GPT-5.6 Sol at up to 750 output tokens per second — 14× faster than Standard mode — powered by Cerebras wafer-scale hardware.
On August 13, 2026, OpenAI unveiled Ultrafast, a new API service tier that runs its flagship GPT-5.6 Sol model at up to 750 output tokens per second — a staggering 14× speedup over Standard processing. Powered by Cerebras Systems’ wafer-scale inference hardware, Ultrafast is being previewed to a select group of API customers and is designed for time-critical, production-grade workloads: real-time voice assistants, live customer support agents, interactive coding tools, and any application where sub-second response latency is the difference between a product that feels magical and one that feels broken.
The launch marks the first time OpenAI has offered a dedicated speed tier for its most intelligent model, and it signals a strategic shift in how the company thinks about the inference market — not just as a question of raw intelligence per token, but of how fast that intelligence can be delivered to end users.
What Ultrafast Delivers
At the core of the announcement is a simple but dramatic claim: 750 output tokens per second. For context, a typical human reads at roughly 4–5 tokens per second. Ultrafast generates text approximately 150 times faster than a person can read it. At that speed, a full page of text — roughly 500 words — can be produced in under a second.
The key specifications:
- Speed: Up to 14× faster than Standard mode processing
- Throughput: Up to 750 output tokens per second
- Model: GPT-5.6 Sol, OpenAI’s flagship frontier model
- Hardware: Cerebras wafer-scale engine (WSE) systems
- Intelligence: No quality degradation — Ultrafast runs the full GPT-5.6 Sol model without distillation or compression
- Availability: Preview access for select API customers, with broader rollout expected
This is not a smaller, faster, dumber model. It is the same GPT-5.6 Sol that OpenAI launched in July 2026 — the model that set state-of-the-art results across coding, knowledge work, and scientific reasoning — now running on hardware optimized to extract maximum inference speed.
The Cerebras Partnership: Why Wafer-Scale Matters
The Ultrafast tier is powered by Cerebras Systems (NASDAQ: CBRS), the chipmaker that has bet its entire existence on a radical idea: instead of fabricating dozens of small chips from a single silicon wafer and cutting them apart, keep the entire wafer intact as a single, massive processor. The result is the Wafer Scale Engine (WSE) — a chip roughly the size of a tablet that packs more compute cores, memory, and bandwidth onto a single die than any GPU on the market.
The architectural advantage is straightforward but profound. In traditional GPU clusters, a large language model’s parameters must be distributed across multiple chips connected by high-speed interconnects (like NVIDIA’s NVLink). Every token generated requires shuttling data back and forth between these chips, creating latency overhead that compounds with model size. Cerebras eliminates this bottleneck by placing the model’s computation on a single wafer, where every core can access shared memory without crossing chip boundaries.
This is why Cerebras can deliver 750 tokens per second on a frontier-class model. The WSE can execute logic across its entire memory space every clock cycle, meaning there is no interconnect tax on inference throughput. The same model that runs at roughly 50–55 tokens per second on standard GPU infrastructure suddenly runs 14 times faster.
The partnership builds on a landmark deal announced in June 2026, in which OpenAI committed to 750 megawatts of Cerebras compute capacity — a deployment valued at over $20 billion over multiple years. At the time, many analysts viewed the deal as a long-term infrastructure hedge. With Ultrafast, that capacity is now being productized into a commercial offering that directly competes on speed with every other inference provider in the market.
The Speed Tier Strategy
Ultrafast sits above two existing processing modes in OpenAI’s API: Standard (baseline) and Fast (launched July 30, delivering up to 2.5× speed at 2× cost). The tier hierarchy now looks like this:
| Tier | Speed vs. Standard | Use Case |
|---|---|---|
| Standard | 1× | Batch processing, non-time-sensitive tasks |
| Fast | Up to 2.5× | General interactive applications |
| Ultrafast | Up to 14× | Real-time voice, live agents, sub-second latency |
This tiering reflects a deeper truth about the AI inference market: speed is becoming a primary axis of competition. For years, the race was about intelligence — whose model scored highest on benchmarks, whose reasoning was most accurate. But as frontier models have converged in capability (GPT-5.6 Sol, Claude Opus 4.6, Gemini 3.1, and DeepSeek V4 all perform within a few percentage points of each other on most benchmarks), differentiation is shifting to latency, throughput, and cost-per-query.
OpenAI’s pricing for the new tier has not been fully detailed in the preview announcement, but based on the Fast mode precedent — which charges 2× Standard pricing for 2.5× speed — Ultrafast is expected to carry a significant premium. Early community estimates suggest the cost could be 18–40× Standard pricing, reflecting the extreme infrastructure requirements of running frontier models at 750 tokens per second. For enterprise customers building real-time products, however, the price premium is secondary to the latency advantage — a voice assistant that responds in 300 milliseconds versus 3 seconds is qualitatively a different product.
Why 750 Tokens Per Second Matters
To understand why this specific number is significant, consider what it unlocks:
Real-time voice AI: Voice assistants require end-to-end latency under 500 milliseconds to feel conversational. At 750 tokens per second, GPT-5.6 Sol can generate a full spoken-length response (roughly 150–200 tokens, or 2–3 sentences) in under 300 milliseconds — well within the window for natural conversational flow. This is the technical foundation for OpenAI’s GPT-Live voice product, launched in July 2026 with SynthID watermarking.
Multi-agent systems: Modern AI applications increasingly involve chains of model calls — an agent that plans, calls tools, iterates on results, and synthesizes a final answer. At Standard speeds, a 5-step agent pipeline might take 15–20 seconds. On Ultrafast, the same pipeline completes in 1–2 seconds, making complex agentic workflows viable for real-time interfaces.
Interactive coding: Developer tools like code completion and automated refactoring require near-instantaneous feedback to maintain flow state. At 750 tokens per second, GPT-5.6 Sol can generate an entire function implementation in the time it takes a developer to lift their finger from the keyboard.
Live customer support: Customer service agents that need to parse a user’s message, consult a knowledge base, and compose a helpful response can now do so in under a second, eliminating the awkward “typing…” delays that have plagued AI-powered support chatbots.
Competitive Implications
The Ultrafast launch puts pressure on every inference provider in the market. Anthropic’s Claude models, Google’s Gemini family, and open-weight alternatives like DeepSeek V4 and Qwen 3.8 all run on GPU-based infrastructure where the interconnect bottleneck caps throughput at roughly 50–80 tokens per second for frontier-class models.
Cerebras itself offers third-party inference — running models from Meta, Mistral, and others on its wafer-scale hardware — and has demonstrated speeds exceeding 2,000 tokens per second for smaller models. But the OpenAI partnership gives Ultrafast a unique positioning: the combination of OpenAI’s most intelligent model with Cerebras’s fastest hardware, sold as a unified API product.
For NVIDIA, the implications are nuanced. NVIDIA GPUs remain the dominant platform for AI training, and Ultrafast does not displace them for model development. But for inference — the market that directly serves end users and generates the most API revenue — Cerebras has now demonstrated a concrete, shipping advantage. If Ultrafast proves popular with enterprise customers, other model providers may seek similar partnerships with Cerebras, Groq, or other specialized inference hardware makers, gradually eroding NVIDIA’s grip on the inference market.
The Road Ahead
Ultrafast is currently in preview, accessible through an interest form on OpenAI’s website. The company has not announced a general availability date or detailed pricing, but the preview’s existence confirms that the Cerebras infrastructure is already operational and serving real traffic.
The broader signal is clear: the AI inference market is entering a new phase where latency is the product. Frontier intelligence is no longer the sole differentiator — it is table stakes. The companies that win the next phase of AI deployment will be those that deliver frontier-grade intelligence at speeds that feel instantaneous to human users. With Ultrafast, OpenAI and Cerebras have drawn first blood in that race.
For developers, the message is direct: if your application has real-time latency requirements — voice, agents, live interaction — there is now a tier built specifically for you. The era of choosing between intelligence and speed is ending. With Ultrafast, you get both.
Sources
- [1] https://openai.com/index/previewing-ultrafast/
- [2] https://techcrunch.com/2026/08/13/openai-introduces-ultrafast-a-new-mode-that-makes-gpt-5-6-sol-work-at-14x-the-speed/
- [3] https://www.cerebras.ai/blog/accelerating-gpt-5-6-sol-ultrafast-with-openai
- [4] https://alphasignal.ai/news/openai-s-gpt-5-6-sol-hits-750-tokens-per-second-on-cerebras-hardware
- [5] https://www.unite.ai/cerebras-runs-openais-gpt-5-6-sol-at-750-tokens-per-second-in-new-ultrafast-tier/
- [6] https://cryptobriefing.com/openai-ultrafast-mode-gpt-5-6-sol/