From Megawatts to Tokens: Nvidia's DSX Posts First Production Numbers — 24% More Throughput on the Same Power
At AI Infra Summit, Nvidia published DSX's first real-world results: Lambda squeezed 24% more token throughput from a fixed power budget, and the Eos AI factory has answered 200+ utility demand signals without interrupting a single critical workload.
On a sweltering August evening in Silicon Valley, as the sun set and air-conditioning loads spiked across Santa Clara, the city’s municipally owned utility did something quietly historic: it sent a demand signal to an AI factory, asking it to consume less electricity. Nobody in the building touched anything. Software read the grid’s condition, demoted the lowest-priority training jobs, kept every high-priority inference workload running, and the facility’s power draw dropped from four megawatts to three — automatically, in under a minute.
That facility is Nvidia’s Eos AI factory, the orchestration layer is Emerald AI’s Conductor platform, and the utility is Silicon Valley Power. Silicon Valley Power has since sent more than 200 demand signals. The system responded correctly every single time.
At the AI Infra Summit in Santa Clara this week — the industry’s largest gathering dedicated purely to AI infrastructure, with more than 8,000 attendees — Nvidia VP of hyperscale and HPC Ian Buck made AI factory efficiency the centerpiece of his keynote, “Advancing Infrastructure for the Era of Agentic AI.” Alongside the keynote, Nvidia published the first production results for its DSX AI factory platform, and for the first time the “AI factory” concept comes with hard, independently-validated numbers attached.
The Lambda validation: 24% more tokens, zero more megawatts
The headline number comes from Lambda, the GPU cloud provider serving more than 10,000 customers from AI-native startups to hyperscalers. Lambda ran DSX MaxLPS — the power-allocation layer of the DSX stack — on a five-rack, 19-node cluster of NVIDIA HGX B200 GPU servers. It is the first validation of MaxLPS on HGX B200 in a real deployment environment.
The mechanics are deceptively simple. Data centers provision power statically: each rack gets a fixed allocation sized for worst-case draw, and the difference between provisioned and consumed power is “stranded” — bought, built, and never used. Training and inference draw power very differently, and in mixed-workload factories the mismatch leaves substantial headroom locked up. DSX MaxLPS monitors GPU- and rack-level consumption in real time and dynamically reallocates that headroom across nodes based on workload type.
The result: Lambda ran 19 nodes inside the same power budget that static provisioning would reserve for 16 nodes at full draw. Cluster-wide token throughput rose 24% — from roughly 4 million tokens per second to 5 million — and performance per watt improved 23%.
“With our proof of concept, we believe we’ve moved beyond the limitation of fixed power budgets,” said Dave Ward, president of cloud services at Lambda. “NVIDIA DSX MaxLPS paves the way to reclaiming stranded capacity and converting it into real-world usage, with significantly more compute density in the same footprint.”
Nvidia’s own projections go further: on next-generation Vera Rubin NVL72 AI factories, the company claims MaxLPS can enable up to 40% more GPU capacity within the same megawatt budget in suitable deployment environments.
Why this matters more than another benchmark
The AI infrastructure conversation of the past two years has been dominated by chip performance — FLOPS, memory bandwidth, interconnect topology. But the binding constraint has shifted. As Jensen Huang has framed it: “A one-gigawatt factory will never become a two-gigawatt factory.” Power delivery, not silicon, is now the scarce resource — grid interconnection queues in major markets stretch for years, and the political tolerance for AI data centers eating city-scale amounts of electricity is thinning fast.
That reframes the optimization problem. If you cannot get more megawatts, the metric that matters is work per megawatt — tokens per second per unit of power. DSX extends the efficiency discipline that operators have long applied at the facility level (cooling, power conversion, rack design) down into the AI workload itself: smarter rack provisioning puts power where workloads actually need it, and operational intelligence — tighter scheduling, faster restarts, leaner checkpointing — keeps GPUs computing rather than waiting.
The Eos/Silicon Valley Power deployment points at the second-order opportunity: grid participation. Nvidia’s Eos factory participates in Silicon Valley Power’s Flexible Load Interconnect Program, described as the first commercial grid-utility program to treat AI factories as dispatchable resources. An AI factory that can credibly flex its demand gets to run bigger — it becomes an asset the grid can work with rather than a load the grid must defend against. Emerald AI’s Conductor is integrating into DSX Flex as the platform matures, and the first dedicated DSX Flex commercial deployment will be a 96-megawatt Vera Rubin AI factory at Nvidia’s AI Factory Research Center in Manassas, Virginia, planned in collaboration with Digital Realty, EPRI and PJM Interconnection.
The full DSX stack
DSX (introduced at GTC Taipei on May 31, 2026) is Nvidia’s answer to a question Huang poses constantly: stop optimizing the parts, start designing the whole. The platform spans:
- DSX Reference Designs — generation-specific validated architectures covering compute, networking, storage and facilities, co-designed with Nvidia’s ecosystem partners
- DSX Sim — simulation tools to model and validate a factory design before capital is committed, catching bottlenecks before the first rack ships (relevant when a GB200 NVL72 rack carries roughly 120 kW of heat to remove via direct liquid cooling)
- DSX OS — open-source, modular software for factory lifecycle management, runtime consistency, health automation and resiliency
- DSX MaxLPS — the power-allocation layer validated by Lambda
- DSX Flex — the grid-signal orchestration layer (load-shedding, demand-response, pricing events) proven in concept by the Eos deployment
- DSX Exchange — secure data exchange across IT and operational-technology systems
New in the September 15 announcement: DSX reference designs are incorporating an 800 VDC power architecture, designed to reduce conversion complexity, improve delivery efficiency and support denser accelerated-computing racks.
The ecosystem is already broad. As of the May launch, CoreWeave, Crusoe, Firmus, IREN, Lambda, Nebius, Nscale and Yotta Data Services were deploying DSX Sim, MaxLPS and OS, while Dell, HPE, Lenovo and Supermicro build DSX-ready systems.
The skeptic’s read
Two caveats are worth holding onto. First, 24% and 23% are Lambda’s numbers for one five-rack cluster in one deployment environment; the 40% Vera Rubin figure is an Nvidia projection “in suitable deployment environments” — a qualifier doing real work. Second, demand-response economics cut both ways: a factory that yields capacity during grid stress needs the revenue or tariff terms to justify curtailed token production. The 200-for-200 success rate at Eos is genuinely impressive, but Eos is a single site with a unusually cooperative municipal utility.
Still, the direction is right and the timing is not accidental. Agentic AI is transforming the demand profile placed on infrastructure — agents tolerate latency in ways batch training does not, which is exactly the flexibility MaxLPS and Flex monetize. The era in which AI infrastructure competed purely on peak performance is giving way to one where the winners are those who extract the most tokens from every watt, and who can shake hands with the grid instead of fighting it. On today’s evidence, Nvidia intends to own that layer too.
Sources
- [1] https://blogs.nvidia.com/blog/from-megawatts-to-tokens-how-nvidia-maximizes-ai-factory-production/
- [2] https://www.unite.ai/nvidia-reports-early-production-results-for-dsx-ai-factory-platform/
- [3] https://www.nvidia.com/en-us/events/ai-infra-summit/
- [4] https://www.nvidia.com/en-us/data-center/products/dsx/