NVIDIA's Groq 3 LPX Enters Full Production: 3,400 Tokens/sec for the Agentic AI Era
NVIDIA's first product from its $20B Groq acqui-hire — the Groq 3 LPX inference rack — is now in full production, pairing 256 LPU accelerators with Vera Rubin NVL72 and posting a record 3,400 output tokens/sec on Gemma 4 31B.
The most consequential silicon story of the Rubin era just moved from slideware to shipping hardware. NVIDIA has announced that Groq 3 LPX, the interactive inference accelerator born from its blockbuster Groq acquisition, is now in full production — and Nebius, the AI cloud spun out of Yandex’s former infrastructure arm, has signed on as the first customer, deploying LPX racks inside its Token Factory inference platform alongside Vera CPUs and Rubin GPUs.
For anyone tracking the AI infrastructure race, this is a watershed moment: it is the first time NVIDIA has fused its rack-scale GPU empire with Groq’s deterministic, SRAM-first LPU architecture into a single commercial platform — and the first hard evidence that the $20B acqui-hire of Jonathan Ross’s team is paying off.
What Groq 3 LPX actually is
NVIDIA positions LPX not as a replacement for its GPUs but as an extension of the Vera Rubin platform. The division of labor is clean:
- Vera Rubin NVL72 remains the versatile training-and-inference workhorse for every AI factory.
- Groq 3 LPX is purpose-built for one thing: ultrafast token generation — the rate at which output tokens stream to an individual user or agent.
Each LPX rack couples 256 LPU accelerators with 128 GB of on-chip SRAM and 640 TB/s of scale-up bandwidth, fully liquid-cooled in NVIDIA’s MGX rack architecture, slotting next to Vera CPU racks, Vera BlueField-4 STX storage, and Spectrum-6 SPX Ethernet. NVIDIA describes the result as “extreme codesign across seven chips and five purpose-built racks” — the most extensive AI factory platform it has ever assembled.
The physics rationale is straightforward. Agentic systems burn tokens asymmetrically: enormous prompt context going in (codebases, tool outputs, multi-turn state), comparatively few tokens coming out — but those output tokens must arrive fast, because agents make dozens or hundreds of sequential model calls where latency compounds across the whole workflow. GPUs excel at batch throughput; LPUs excel at deterministic, low-latency generation. Rubin plus LPX covers both ends.
The numbers
In Artificial Analysis benchmarking, Groq 3 LPX delivered a record 3,400 output tokens per second running Gemma 4 31B — an open-source agentic model — at a 100,000-token context, the fastest performance ever recorded for that model. NVIDIA claims:
- 4x faster responsiveness for agents and latency-sensitive workloads versus the nearest alternative platform
- Agentic coding tasks completed in minutes instead of hours
- Up to 35x higher inference throughput per megawatt for 2T-parameter models at long context and low latency compared to GB200 NVL72, per NVIDIA’s own projections
That per-megawatt figure deserves emphasis. At a moment when AI factories are power-constrained before they are chip-constrained — and when NVIDIA has just warned hyperscalers of 15%+ price hikes on Rubin and Blackwell systems driven by soaring DRAM costs — an inference rack that leans on on-chip SRAM instead of HBM is a strategically elegant hedge.
Nebius first through the door
Nebius Token Factory becomes the first production home for LPX, and the company’s framing is telling. “Generation is the phase of inference that determines how responsive an AI system actually is, and that’s exactly what NVIDIA Groq 3 LPX is built to accelerate,” said Nebius CTO Danila Shtan. “As the first AI cloud bringing it to production via Nebius Token Factory, we’re making sure every step of an agent’s loop feels instant — through the same API developers are already using, with no migration to a new stack.”
That last clause is the quiet strategic point. A fast token means little if adopting it requires re-architecting your stack for exotic silicon — the trap Groq the startup repeatedly fell into. Inside Token Factory, LPX is exposed as a model-selection change, not a migration: same API, same autoscaling and observability, same billing — initially for a subset of models, with native function calling, structured JSON outputs, and safety guardrails for tool-calling agents intact.
Following Nebius, NVIDIA says a purpose-built AI inference cloud plans to be among the platform’s earliest adopters.
Why this matters
Three implications stand out.
1. The acqui-hire thesis is validated. When NVIDIA absorbed Groq — the LPU pioneer founded by a Google TPU architect — skeptics asked whether the architecture would survive inside a GPU-centric giant or be shelved. Full production within the Rubin timeline answers emphatically: the LPU lives on as a first-class NVIDIA product line, with the “Groq” name and LPU trademarks used under license from Groq, Inc.
2. Inference is splitting into two markets. Jensen Huang calls inference “the growth engine of AI,” and the LPX launch formalizes a split that has been visible for a while: bulk batch inference (where GPU throughput per watt wins) versus interactive agentic inference (where single-stream token speed wins). Expect pricing, benchmarks, and procurement conversations to bifurcate accordingly.
3. The speed bar just moved. A coding agent that can iterate at thousands of tokens per second changes what “agent time” costs. If agents complete in minutes what previously took hours, the economics of multi-agent pipelines — and the token budgets they consume — shift again in favor of more autonomous software.
Caveats worth noting: there’s no public API availability date yet, the 35x/MW figure is NVIDIA’s projection for a specific 2T-parameter long-context regime rather than an independent measurement, and LPX supports only a subset of models at launch. The record-setting Gemma 4 31B result, while real, is a 31B model — frontier-scale LPX performance remains to be demonstrated.
Even so, the direction is unmistakable. With SpaceX also disclosed as adopting NVIDIA Vera CPUs for its next-generation AI stack — including orbital data centers — and hyperscalers preparing for Rubin shipments in early 2027, NVIDIA is methodically assembling every piece of the agentic AI compute stack. Groq 3 LPX in full production is the piece nobody was sure would survive — and it just shipped.
Sources
- [1] https://nvidianews.nvidia.com/news/nvidia-groq-3-lpx-now-in-full-production-with-world-class-speed-for-agentic-ai
- [2] https://nebius.com/blog/posts/nvidia-groq-3-lpx-nebius-token-factory
- [3] https://www.hpcwire.com/bigdatawire/this-just-in/nvidia-groq-3-lpx-enters-full-production-for-agentic-ai-inference/
- [4] https://www.techmeme.com/260824/p21