NVIDIA's Groq 3 LPX Enters Full Production: 3,400 Tokens/Sec Inference Racks Built for the Agentic Era
NVIDIA's Groq 3 LPX — the low-latency inference rack born from its $20B Groq acqui-hire — is now in full production, pairing 256 LPU accelerators with 128GB of on-chip SRAM per rack and hitting a record 3,400 output tokens/sec on Gemma 4 31B. Nebius is the first AI cloud to deploy it inside its Token Factory.
At Hot Chips in Palo Alto this week, NVIDIA announced that its Groq 3 LPX — the rack-scale interactive AI inference accelerator built out of its blockbuster $20 billion Groq acqui-hire — has entered full production, with Nebius signed as the first AI cloud customer. The launch marks the most consequential reshaping of the AI inference market since Groq’s LPU architecture first made its name on breakneck decode speeds: the startup’s deterministic, SRAM-heavy silicon philosophy is now shipping as an official extension of NVIDIA’s Vera Rubin platform, the successor to the Blackwell NVL72 generation that powers most frontier AI factories today.
What Groq 3 LPX actually is
Groq 3 LPX is not another GPU. It is a purpose-built inference rack designed to slot alongside NVIDIA Vera Rubin NVL72 systems and attack one specific bottleneck: token generation latency.
Agentic AI has created an unusual computing problem. An agent working through a complex task might make dozens or hundreds of sequential model calls — reading files, writing and testing code, calling tools, verifying results — and every step depends on the one before it. Latency doesn’t just add up across that chain; it multiplies into the user experience. A model that reasons brilliantly but decodes slowly makes for an agent that feels sluggish and misses real-time deadlines.
NVIDIA’s answer is to split inference into its two natural jobs. Vera Rubin’s GPU racks handle context processing — ingesting long prompts, large codebases, and persistent multi-turn agent state. The LPX rack, built from LPU (Language Processing Unit) accelerators, handles the decode phase, where tokens are generated one at a time and raw clock speed matters most.
The rack-level numbers are striking. Each Groq 3 LPX rack couples 256 LP30 accelerators with 128 GB of aggregate on-chip SRAM — the same memory technology that made Groq’s original chips legendary for latency, since SRAM sidesteps the HBM bandwidth ceiling that constrains GPU decode. NVIDIA’s spec sheet lists 40 petabytes per second of aggregate SRAM bandwidth per rack and 640 TB/s of scale-up bandwidth linking the accelerators, all fully liquid-cooled inside the NVIDIA MGX rack architecture. At the individual chip level, each LP30 accelerator carries 500 MB of on-chip SRAM with 150 TB/s of bandwidth.
The benchmark that matters: 3,400 tokens per second
In independent benchmarking by Artificial Analysis, Groq 3 LPX running Gemma 4 31B — an open-source agentic model — delivered a record 3,400 output tokens per second for a single user at a 100,000-token context length. NVIDIA says that is the fastest performance ever recorded for that model, and roughly 4x faster responsiveness for latency-sensitive workloads than the nearest alternative platform.
To put that in perspective: a frontier chat model typically streams 40–100 tokens per second to a human reader. Agentic coding systems can consume thousands of tokens per second when they’re iterating autonomously. At 3,400 tokens per second, an agent can finish generation phases in seconds that would otherwise take minutes — which NVIDIA frames as enabling “coding in minutes versus hours.”
The efficiency claim is equally aggressive. NVIDIA projects that Vera Rubin NVL72 combined with Groq 3 LPX can deliver up to 35x higher inference throughput per megawatt for 2-trillion-parameter models in long-context, low-latency configurations compared to GB200 NVL72. Whether that projection survives contact with real-world workloads remains to be seen, but the direction is clear: the metric that AI factories now optimize for is tokens per second per megawatt, not raw FLOPS.
Nebius first, Groq itself next
Nebius, the AI cloud spun out of Yandex, will be the first to deploy LPX racks in production through its Token Factory inference platform. Notably, Nebius is positioning the integration as a non-event for developers: LPX-accelerated models will run behind the same API, with the same autoscaling, observability, function calling, and structured outputs — “a model-selection change, not a migration,” as the company put it.
“Generation is the phase of inference that determines how responsive an AI system actually is, and that’s exactly what NVIDIA Groq 3 LPX is built to accelerate,” said Danila Shtan, CTO of Nebius. “As the first AI cloud bringing it to production via Nebius Token Factory, we’re making sure every step of an agent’s loop feels instant.”
In a twist that would have seemed implausible two years ago, Groq itself — now an NVIDIA subsidiary operating its own inference cloud — plans to be among the platform’s earliest adopters after Nebius. The once-fierce rival has become both supplier and customer.
Analysis: inference is the new battlefield
The strategic read here is straightforward: NVIDIA is segmenting the AI factory market precisely as agentic workloads explode. Training monopolized the last two hardware cycles; the next one belongs to inference, where token economics decide who profits. Jensen Huang’s statement accompanying the launch said it plainly: “Inference is the growth engine of AI… Vera Rubin extends that vision with workload-optimized AI factory configurations designed for the era of agentic AI.”
The full Vera Rubin platform NVIDIA detailed at Hot Chips is an exercise in what the company calls “extreme codesign” — seven chips and five purpose-built rack types, spanning Vera CPU racks for the orchestration-heavy parts of agent loops, BlueField-4 DPUs, Vera BlueField-4 STX storage, Spectrum-6 SPX Ethernet, and now the LPX inference rack. The same event brought news that CoreWeave has put Spectrum-X Multiplane into production connecting Vera Rubin racks, and that SpaceXAI will build its next-generation agentic stack — including its planned orbital data centers — on Vera CPUs.
For buyers, one caution tempers the enthusiasm: this launch lands days after reports that NVIDIA has warned hyperscalers of 15%+ price increases on Rubin and Blackwell systems starting in early 2027, driven by surging memory costs that even 75% gross margins can no longer absorb. Cutting-edge inference throughput is arriving — but the bill for it is rising too.
What happens next is worth watching closely. If LPX-class decode performance becomes the default for agentic serving, the competitive pressure on GPU-only inference incumbents will intensify, and the “tokens per second per megawatt” metric could become the industry’s defining benchmark — the inference-era equivalent of the training-FLOPS race that defined the last three years of AI hardware.
Sources
- [1] https://nvidianews.nvidia.com/news/nvidia-groq-3-lpx-now-in-full-production-with-world-class-speed-for-agentic-ai
- [2] https://blogs.nvidia.com/blog/vera-rubin-lpx-spectrum-x-nvlink-fusion/
- [3] https://www.nvidia.com/en-us/data-center/lpx/
- [4] https://nebius.com/blog/posts/nvidia-groq-3-lpx-nebius-token-factory
- [5] https://wccftech.com/nvidia-groq-3-lpx-ai-inference-accelerator-full-production-supercharging-vera-rubin/