7x Per Megawatt: SemiAnalysis Catches Nvidia Sandbagging Vera Rubin's Real Inference Power
SemiAnalysis's first verified AgentX results for Vera Rubin NVL72 show up to 7x better token throughput per megawatt than Blackwell on DeepSeek V4 Pro — more than double Jensen Huang's 3x GTC claim — worth ~$149.9B in modeled annual profit per gigawatt.
Every Nvidia product cycle has a ritual. Jensen Huang takes the GTC stage, shows a bar chart, and the industry spends the next year discovering the chart was conservative. With Vera Rubin, the discovery phase just arrived early — and the gap between claim and reality is bigger than anyone outside SemiAnalysis’s lab expected.
On September 14, the research firm published the first independently verified agentic inference results for Nvidia’s next-generation Vera Rubin NVL72 rack, measured on its AgentX benchmark, which replays real-world agentic traffic across a fleet of thousands of chips. The headline finding is blunt enough to become an industry talking point: on the DeepSeek V4 Pro workload (a 1.6-trillion-parameter MoE model that serves as a proxy for the O(1–3T) frontier), Vera Rubin NVL72 on early pre-release software already delivers up to 7x better token throughput per megawatt than Blackwell — against the 3x figure Jensen Huang himself presented at GTC 2026.
SemiAnalysis’s verdict, in its own words: “Jensen needs to stop sandbagging his performance claims at GTC.”
Why ‘sandbagging’ is the operative word
This is not an accusation of deception — it is an accusation of radical understatement, and SemiAnalysis has receipts from the last cycle. At GTC 2024, Huang claimed GB200 NVL72 would deliver 30x Hopper’s performance. When SemiAnalysis eventually ran the hardware through InferenceX, it measured 98x. The pattern the firm is documenting is that Nvidia’s public generational claims systematically undersell what the shipping platform does once the software stack matures — which means buyers and competitors who model their procurement on GTC slides are working from numbers that are, in a precise sense, wrong by multiples.
The power-efficiency numbers deserve the attention they are getting. At a 100 tokens-per-second interactivity target on the DeepSeek V4 Pro agentic workload, Vera Rubin delivers approximately 59.4 million total tokens/sec per megawatt, compared with 28.5 million for GB300 running Dynamo SGLang (the stronger of the two GB300 engines at that target) and 21.1 million for GB300 Dynamo TRTLLM. That’s a 2.09x advantage at 100 TPS — but the curve steepens dramatically with interactivity. At 150 TPS, Rubin retains ~37 million tok/s/MW, roughly 7.2x GB300 SGLang. At exactly 170 TPS, Rubin delivers 62.9x the throughput per MW of GB300 TRTLLM at its measured endpoint — though SemiAnalysis is careful to note the engine label matters enormously: against GB300 SGLang at 170 TPS, the multiplier is 5.56x.
For everyone not running a GPU fleet, the comparison table is grim for the competition: MI355X SGLang reaches 2.01 million tok/s/MW at 100 TPS (a 29.5x Rubin lead), B200 SGLang 6.95 million, B300 vLLM 5.56 million, and H200 Dynamo SGLang 2.26 million on FP8. On agentic workloads, SemiAnalysis writes, “Vera Rubin makes H200 look about as competitive as a TI-84 calculator” — 18x more tokens per dollar at 80 TPS P90, widening to 39x at 120 TPS.
The economics: $149.9 billion of profit per gigawatt
The reason per-megawatt numbers dominate this cycle is that power, not chips, is the binding constraint on AI buildouts. Availability of powered data centers is the limiting factor to deploying more accelerators, so tokens-per-gigawatt translates directly into revenue and profit capacity without securing new utility contracts.
SemiAnalysis modeled it. At 75 TPS interactivity, 60% utilization, and no model-license fee (DeepSeek V4 Pro ships under MIT), Vera Rubin NVL72 generates $159.5 billion in annual revenue and $149.9 billion in modeled profit per all-in utility gigawatt. The strongest GB300 configuration (Dynamo SGLang) generates $114.9 billion and $105.3 billion respectively — meaning Rubin yields roughly 39% more revenue and 42% more profit from the same power allocation, an absolute difference of about $44.6 billion in additional annual modeled profit per GW.
That efficiency also creates pricing headroom: holding workload mix constant, a Rubin operator could charge approximately 28% less per token across cached-input, uncached-input, and output pricing while still matching GB300’s revenue per gigawatt. In a market where DeepSeek-class open-weight models have already dragged frontier prices down, that is a strategic weapon — an operator can bank the efficiency as margin or deploy it to undercut rivals, depending on demand elasticity.
On total cost of ownership, the story repeats. Using the “owning at hyperscaler volume” cost model at 170 TPS on apples-to-apples TRTLLM NVFP4 dense configs, Vera Rubin delivers ~67x the total throughput per TCO dollar of GB300 Dynamo TRTLLM. In the 60–100 TPS band where providers would actually serve this model, Rubin achieves 1.4x–3x the throughput per TCO versus the latest GB300 TRTLLM. Even on 3-year rental economics — where Rubin rented at over $8.5/hr/chip in July 2026 against $5/hr for Blackwell Ultra NVL72 — the upgrade pays: at 80 TPS, Rubin produces 62% more total tokens for the same rental TCO, and up to 16x more tokens per rental dollar at higher interactivity targets.
Co-design, KV cache, and the agentic workload
The article’s technical core explains why the gaps are so large: Rubin is the first platform co-designed across six products specifically for the agentic era — the Rubin GPU, Vera CPU, NVLink 6 switch, ConnectX-9, BlueField-4, and Spectrum-6. AgentX’s workload model reflects how agents actually run: multi-turn sessions with tens or hundreds of interactions, long accumulating contexts, high prefix reuse (where turn n’s context is mostly served from KV cache rather than recomputed), and bursty sub-agent fan-outs that create spiky KV-cache pressure. Chatbot-era benchmarks barely touch this profile.
A second efficiency lever is new to this generation: DSX MaxLPS, a first-class dynamic power-shifting integration. Instead of provisioning clusters for maximum TDP plus a 10–20% oversubscription factor, operators can power-profile inference workloads and steer power across the data center — fitting more GPUs into the same power footprint, since GPUs at medium-to-fast serving speeds rarely consume their full envelope. SemiAnalysis says its upcoming PowerX integration into InferenceX will measure throughput-per-gigawatt at even finer grain.
There are honest caveats in the piece. The Rubin numbers come from an early pre-release TRTLLM build — performance “will only get better,” especially at the extremes of the latency-throughput frontier. Vera Rubin hits ~61% higher maximum P90 interactivity than GB300 TRTLLM (276 vs 171 P90 TPS), though the open-source SGLang stack narrows that particular gap. And when GB300 uses SGLang instead of TRTLLM, some interactivity comparisons converge. The tested rack uses the production SKU: 2300W TDP per compute tray, 1.5TB of CPU LPDDR5X.
What it means
Three implications follow from the data.
For hyperscalers and neoclouds, the article’s takeaway is unambiguous: “If you have the money to buy or rent a VR NVL72, you should — it will generate significantly cheaper tokens than the next leading accelerator and thus make you significantly more money.” The more you buy, the more you earn — Huang’s infamous line — is, on SemiAnalysis’s math, now quantifiably true at the gigawatt scale.
For AMD, the 29.5x per-megawatt gap to MI355X on this workload snapshot is the kind of number that ends conversations in procurement meetings, though SemiAnalysis notes AMD has committed to collaborating on MI455X UALink72 with InferenceX — the benchmark is the only one running TPU v7, Nvidia, AMD, and soon SambaNova and Trainium.
For everyone modeling the AI buildout, the meta-lesson is the sandbagging pattern itself. GTC keynote numbers are the floor, not the ceiling, of what a shipping Nvidia platform delivers on agentic workloads. Anyone constructing TCO models, competitive analyses, or sovereign-AI procurement plans from Jensen’s bar charts should apply a historical correction factor — in the direction of Nvidia being even further ahead than advertised.
The publishing context matters too: this analysis landed the same week Nvidia entered the Dow-fresh aftermath of a $96.2 billion quarter with Vera Rubin ramping to full production across CoreWeave, Azure, Google Cloud, and Oracle — and the same week the industry’s attention was consumed by AI-safety resignations and pacing debates. The quiet story of the week may turn out to be that the hardware frontier just moved again, by more than anyone was told.
Figures and quotes from the SemiAnalysis AgentX/InferenceX report “Vera Rubin NVL72 Agentic Inference: 67x better Performance per Dollar” (Sep 14, 2026, partially paywalled).
Sources
- [1] https://newsletter.semianalysis.com/p/vera-rubin-nvl72-agentic-inference
- [2] https://www.techmeme.com/
- [3] https://www.coreweave.com/blog/nvidia-vera-rubin-nvl72-on-coreweave-10x-more-tokens-per-megawatt-than-blackwell
- [4] https://developer.nvidia.com/blog/nvidia-vera-rubin-and-blackwell-set-a-new-standard-for-agentic-ai-performance-per-watt/