← All posts / Industry

NVIDIA's Groq 3 LPX Enters Full Production: 3,400 Tokens/sec and the Coming Disaggregation of AI Inference

At Hot Chips 2026, NVIDIA announced its Groq 3 LPX inference rack has entered full production, delivering a record 3,400 output tokens/sec on Gemma 4 31B at 100K context — with Nebius as the first cloud customer for its Token Factory and SpaceXAI standardizing on Vera CPUs.

NVIDIA's Groq 3 LPX Enters Full Production: 3,400 Tokens/sec and the Coming Disaggregation of AI Inference

NVIDIA has turned the page on the AI inference market. At Hot Chips 2026 in Palo Alto this week, the chipmaker confirmed that its Groq 3 LPX — the rack-scale, low-latency inference accelerator born from its $20 billion Groq acquisition — has entered full production, with neocloud provider Nebius signed on as the first cloud customer. An independent benchmark from Artificial Analysis clocks the system at 3,400 output tokens per second running the open-source Gemma 4 31B model at a 100,000-token context window, which NVIDIA says makes it roughly 4x faster than the nearest alternative platform for latency-sensitive workloads.

For an industry whose economics are shifting from training to serving, this is not a minor speed bump. It is the clearest signal yet that inference is becoming a disaggregated, specialized business — and that NVIDIA intends to own every layer of it.

What NVIDIA actually announced

The Groq 3 LPX is a purpose-built extension of NVIDIA’s flagship Vera Rubin data center platform. It is not a GPU. It is built from LP30 LPU accelerators — deterministic, SRAM-heavy processors descended from Groq’s original Language Processing Unit design — with up to 256 LPUs per rack, linked by ultra-high-bandwidth chip interconnects.

The architecture’s defining feature is on-chip memory. Each Groq 3 LPU carries 500 MB of SRAM, giving a full rack 128 GB of SRAM total with a staggering 40 PB/s of aggregate on-chip bandwidth — versus roughly 1.6 PB/s for a comparable HBM-based GPU system. A 70B-parameter FP8 model can sit entirely inside a single rack’s SRAM, which is precisely why decode — the phase of inference where a model generates tokens one at a time — runs so fast. Total rack-level AI inference compute is rated at 315 PFLOPS, and NVIDIA’s GTC materials pegged system-level inference compute at 8-bit precision around 1.2 PFLOPS per LPU.

The system is built as 32 liquid-cooled 1U compute trays, each integrating eight Groq 3 LPU accelerators alongside a host CPU, and it slots directly into the Vera Rubin platform family — pairing with the Vera Rubin NVL72 rack that handles large-scale context processing with Rubin GPUs and Vera CPUs.

Why decode latency is the new bottleneck

To understand why NVIDIA paid $20 billion for Groq in December — bringing founder Jonathan Ross and President Sunny Madra in-house — you have to understand how agentic AI changed inference.

Agents don’t just answer a question. They reason, plan, write and execute code, inspect system files, and call third-party tools in continuous loops, often across tens of thousands of tokens of context. Each step generates a response one token at a time. Tiny delays in that decode phase multiply across every step of the chain: a five-second-per-step stall becomes an hour-long task across a complex workflow. NVIDIA’s framing is that multistep agentic tasks that took hours can now complete in minutes.

The LPX’s answer is workload disaggregation. Rubin GPUs handle the prefill — ingesting and processing enormous context windows — while LPX accelerators offload the latency-sensitive decode. The two work in tandem as a unified inference engine, designed to eliminate the classic tradeoff between throughput and response time.

Nebius first, and a SpaceXAI endorsement

Nebius Group is the first AI cloud to commit to the platform. It will deploy Groq 3 LPX in the Nebius Token Factory, its production inference platform, alongside Vera CPUs and Rubin GPUs.

“Generation is the phase of inference that determines how responsive an AI system actually is, and that’s exactly what Groq 3 LPX is built to accelerate,” said Nebius CTO Danila Shtan. “As the first AI cloud to bring it to production via Nebius Token Factory, we’re making sure every step of an agent’s loop feels instant.”

The endorsement matters because Nebius — a neocloud that rents AI compute to developers — is exactly the customer class that lives or dies on token economics. If LPX racks deliver the promised responsiveness, expect other inference-first clouds to follow quickly.

NVIDIA also landed a flagship customer announcement of a different kind: SpaceXAI said it will build its next-generation AI architecture around the Vera Rubin platform, deploying Vera CPUs for the CPU-intensive side of agentic AI — orchestration, tool use, code execution, data processing, and simulation — across both terrestrial data centers and orbital satellites. CoreWeave, meanwhile, has put NVIDIA’s Spectrum-X Multiplane Ethernet architecture into production, connecting Vera Rubin racks in parallel switch paths to scale AI clusters toward 512,000 GPUs.

The economics: token factories and the $45-per-million question

NVIDIA’s broader pitch at Hot Chips is the “token factory” — infrastructure engineered to turn compute into intelligence (and revenue) at the lowest cost per token. The company’s earlier GTC materials claimed that pairing a Groq 3 LPX rack with a Rubin NVL72 system could generate a million tokens for $45 on a 1-trillion-parameter GPT-scale model. Combine that with this week’s news that NVIDIA has warned hyperscalers of 15%+ price hikes on Rubin and Blackwell systems starting early 2027 — driven by DRAM costs it says it can no longer absorb even at 75% gross margin — and the strategic picture snaps into focus: NVIDIA is offering differentiated silicon for the decode workload while raising prices on the general-purpose fleet.

That is a two-front strategy competitors will have to answer. Cerebras is pushing wafer-scale CS-4 systems with disaggregated inference claims of its own, and AMD, AWS Trainium, and Google TPU all target the same inference shift. NVIDIA’s bet is that codesigning compute, networking, and inference acceleration as one system — Spectrum-X for data movement, Rubin for context, LPX for generation — keeps the whole factory, not just one chip, inside its ecosystem.

What to watch

Three things will determine whether Groq 3 LPX reshapes the market or remains a premium niche. First, availability: full production is announced, but customer deployment timing through neoclouds like Nebius is the real test. Second, pricing pass-through: with 15%+ system price hikes looming for 2027 shipments, whether LPX’s token economics survive the sticker shock will decide adoption. Third, competitor response: Cerebras, AMD, and the custom-silicon crowd now have a clear target — 3,400 tokens/sec at 100K context.

What is no longer in question is the direction. Inference is splitting into prefill and decode, and the decode market belongs to whoever can make agents feel instant. As Jensen Huang put it: “We’re advancing the performance frontier with LPX for ultra-fast token generation. This transforms how intelligence is produced.”

Sources