← All posts / Models

Cerebras Unveils CS-4: A 750-PFLOP Wafer-Scale Rack That Claims Up to 30x GPU Inference Speed

Cerebras's new rack-scale CS-4, built from three WSE-3 Turbo wafers, delivers 750 PFLOPs, 129.6 PB/s of memory bandwidth, and over 4,400 tokens/sec per user on GPT-OSS-120B — up to 30x faster than GPU systems, with first shipments this quarter.

Cerebras Unveils CS-4: A 750-PFLOP Wafer-Scale Rack That Claims Up to 30x GPU Inference Speed

At 8 PM ET on August 18, 2026, Cerebras Systems (NASDAQ: CBRS) pulled back the curtain on the CS-4 — a rack-scale AI accelerator the company bills as the fastest in the industry, and its most aggressive answer yet to the GPU status quo that Nvidia has dominated for the better part of a decade.

The headline numbers are blunt. On GPT-OSS-120B, with identical prompts, Cerebras says the CS-4 delivers more than 4,400 tokens per second per user — up to 30 times faster than GPU-based solutions. The system packs 750 PFLOPs of AI compute per rack, 7.2 terabits per second of I/O, and 129.6 petabytes per second of aggregate memory bandwidth. First shipments begin this quarter.

Three wafers, one rack

The CS-4 is built from three newly announced Wafer Scale Engine 3 Turbo (WSE-3T) processors — each still the largest AI chip ever made, with four trillion transistors and 900,000 AI-optimized cores spread across 46,225 square millimeters of silicon, plus 44 GB of SRAM integrated directly on the wafer.

The “Turbo” is not marketing gloss. Compared with the standard WSE-3 that powered the CS-3, the WSE-3T doubles AI compute to 250 PFLOPS per wafer and doubles memory bandwidth to 43.2 petabytes per second. On-chip fabric bandwidth jumps to 53.5 PB/s and off-chip I/O to 2.4 Tbps per wafer, while I/O latency shrinks from five microseconds to as low as two. Since memory bandwidth — not raw FLOPs — is usually the binding constraint on inference speed, doubling it is the change that actually moves tokens per second.

It is also the first product built on what Cerebras is calling the Nexus Platform Architecture, a modular rack design split into three independently scalable elements: compute, power, and I/O.

The “Wafer-Scale Backpack”

The most striking engineering choice is the re-imagined compute subsystem. Each wafer now lives in a rear-mounted “Backpack” — a self-contained assembly that folds power conversion, direct liquid cooling, high-speed I/O, and control electronics into a compact three-dimensional package built directly around the wafer.

Two claims stand out. First, deployment time drops “from days to hours,” because the backpack attaches vertically to the power array and decouples compute from the power supplies. Second, the Backpack has 50% fewer components than the prior-generation system and uses 60% more automated manufacturing — a quiet nod to the fact that at wafer scale, manufacturability and field serviceability matter as much as peak specs.

Power delivery got its own overhaul. By moving power conversion 100x closer to the processor — from roughly 50 millimeters away on conventional GPU boards to about 0.5 millimeters — the CS-4 nearly eliminates board-level power loss and delivers twice as much power to the WSE-3T. That, in turn, enables higher operating frequencies and faster token generation. Cerebras also claims up to 10x more throughput per watt than the CS-3, a number that matters enormously in a year when power availability, not chip supply, is the binding constraint on data center buildouts.

The new programmable I/O subsystem supports two connectivity modes. The standards-based path speaks RoCE v2 RDMA over Ethernet, so CS-4 racks can drop into existing infrastructure and heterogeneous clusters. The more interesting path is Direct Wafer Links, which connects wafers within and across racks without a switch, achieving wafer-to-wafer latency as low as two microseconds.

That latency figure is the enabler behind the CS-4’s most forward-looking claim: the ability to build very large clusters and support models with more than 50 trillion parameters. For context, the largest frontier models today are an order of magnitude smaller. Cerebras is explicitly positioning its fabric for a world where model sizes keep climbing and cross-chip communication becomes the bottleneck that breaks GPU clusters first.

The programmable I/O is also designed for heterogeneous disaggregated inference — where a purpose-built prefill engine (from, say, AMD Helios or AWS Trainium, both named as ecosystem partners) processes an incoming prompt and hands decode off to Cerebras for ultra-low-latency token generation.

Speed as strategy

“In AI, speed is productivity,” said Andrew Feldman, CEO and co-founder of Cerebras. “Historically, fast inference meant using smaller and less capable models. Cerebras CS-4 delivers industry-leading speeds on the largest frontier models, fundamentally changing the paradigm.”

CTO Sean Lie made the agentic case more concretely: “Being 30 times faster doesn’t just make a response feel fast. It gives an agentic system room for more than an order of magnitude as much reasoning, verification, or tool use in the same wall-clock time.”

That framing targets the fastest-growing workload in AI. Agents that reason in loops, call tools, and verify their own outputs burn hundreds of thousands of tokens per user task. At 4,400 tokens per second, a one-million-token reasoning trace completes in under four minutes; on a conventional GPU serving stack, the same trace can take over an hour. The speed argument is really a cost-per-completed-task argument — and Cerebras claims the CS-4’s per-watt throughput gains let data centers serve “higher-value tokens and more total tokens within a given power budget.”

SemiAnalysis founder Dylan Patel, quoted in the release, framed CS-4 as the vehicle for scaling “ultrafast tokens for larger models and significant user volumes.”

Context: a public company racing Nvidia’s roadmap

The launch comes exactly a week after Cerebras-powered “Ultrafast” mode brought OpenAI’s GPT-5.6 Sol to 750 tokens per second — the partnership that transformed Cerebras from an exotic research-vendor into infrastructure for one of the world’s most-used AI products. The CS-4 is the hardware designed to make that class of performance repeatable at frontier scale, and its forward-looking-statements section still names OpenAI, G42, MBZUAI, and AWS as the customers its business depends on.

The competitive backdrop is unforgiving. Nvidia’s Rubin generation is expected to raise GPU inference throughput substantially, and Groq, SambaNova, and AMD’s Helios racks are all chasing the same low-latency decode niche. Cerebras’s differentiation remains architectural: one giant wafer avoids the inter-chip tax that GPU clusters pay on every tensor parallel across dozens of H100s or B200s.

Whether “up to 30x” survives independent benchmarking — the footnote in the release itself concedes that “actual throughput varies by model architecture, context length, precision, and serving configuration” — is the question early customers will answer over the next two quarters. But the direction is unambiguous: inference speed has become a first-class product feature, and Cerebras just raised the bar for everyone building token factories in 2026.