← All posts / Models

NVIDIA's Nemotron 3.5 Lightning: A 30B MoE Model That Thinks Like a 3B Model

NVIDIA releases Nemotron 3.5 Lightning — an open 30B mixture-of-experts model with only 3B active parameters, cutting agent inference costs by 58% and runtime by 33% while running 4x faster than comparable dense models.

NVIDIA's Nemotron 3.5 Lightning: A 30B MoE Model That Thinks Like a 3B Model

What NVIDIA Just Shipped

On August 11, 2026, NVIDIA released Nemotron 3.5 Lightning, an open-source AI model that distills the power of a 30-billion-parameter model into the running cost of one ten times smaller. The release targets what has become the most expensive problem in enterprise AI: the cost of always-on autonomous agents that make thousands of inference calls per task.

The core engineering insight is the mixture-of-experts (MoE) architecture. While the model contains approximately 30 billion total parameters, only about 3 billion are activated during any single forward pass. This means the computational cost per generated token is comparable to a dense 3-billion-parameter model, while the effective knowledge capacity and reasoning quality benefit from the full 30-billion-parameter parameter budget. NVIDIA reports the model delivers up to 4x faster output compared to similarly-sized dense models.

The model is free to download, use, and modify — released under NVIDIA’s open model license — and is immediately available across a broad ecosystem of inference providers including OpenRouter, DeepInfra, GMI Cloud, Amazon SageMaker JumpStart, and FriendliAI.

The Numbers That Matter

NVIDIA’s announcement is anchored by real deployment benchmarks rather than synthetic test scores:

  • 58% cost reduction: In Ramp’s SWE-Bench evaluation, using Nemotron 3.5 Lightning via NeMo Switchyard to handle routine agent steps while routing only complex decisions to a frontier model cut total inference spending by 58%.
  • 33% faster runtime: The same Ramp benchmark showed end-to-end task completion time dropping by a third.
  • 4x output speed: Compared to dense models in the same parameter class.
  • $0.10 per million input tokens: Current pricing on OpenRouter, with output tokens at $0.25 per million — among the cheapest in its tier.

Thoughtworks, in an independent evaluation published the same day, confirmed a drop from $0.477 to $0.250 per million output tokens when migrating a worked example from a comparable dense model to Nemotron 3.5 Lightning on stock vLLM. That is a 48% cost saving with no architecture changes required.

NeMo Switchyard: The Router That Makes It Work

Alongside the model, NVIDIA open-sourced NeMo Switchyard, a model routing library that the company describes as the connective tissue for multi-model agent architectures. The idea is straightforward but powerful: not every step in an agent workflow needs a frontier model. Switchyard dynamically routes each inference step to the most cost-effective model capable of handling it — Nemotron 3.5 Lightning for the high-volume, specialized calls, and a frontier model for the genuinely hard reasoning.

This is a strategic move. NVIDIA is not just shipping a model; it is shipping the infrastructure layer that decides which model runs when. As the Futurum Group noted in their day-one analysis, “Nemotron 3.5 Lightning pairs a 30B open MoE model with NeMo Switchyard, an open source router reshaping agentic AI inference.” The implication is that NVIDIA wants to be the default control plane for multi-model AI deployments, not just the GPU vendor underneath them.

Switchyard integrates with NVIDIA’s broader DGX Cloud and RTX workstation ecosystem, meaning the same routing logic can run on edge devices, local workstations, and cloud infrastructure — a unified model-serving stack from the desktop to the data center.

Distilled From a Frontier Model

Nemotron 3.5 Lightning is not a model trained from scratch in isolation. According to GMI Cloud’s technical breakdown, the model was distilled from NVIDIA’s frontier Nemotron 3 Ultra — the company’s top-tier reasoning model released in June 2026. This means Lightning inherits knowledge representations and reasoning patterns from a far larger system, then compresses them into the efficient MoE architecture.

The model architecture uses an interleaved hybrid design combining Mamba-2 state space layers with attention mechanisms. This hybrid approach is increasingly favored for agent workloads because it handles long-context scenarios more efficiently than pure attention models while maintaining the reasoning quality needed for tool-use and multi-step planning.

Why This Matters for Enterprise AI

The economics of agentic AI have been broken for months. A single complex agent task — say, researching a market, summarizing findings, and generating a report — can involve dozens to hundreds of model calls. When each call hits a frontier model costing $5-15 per million tokens, the per-task cost quickly becomes prohibitive at scale. Companies building production agent systems have been forced into painful trade-offs between quality and cost.

Nemotron 3.5 Lightning directly addresses this bottleneck. By providing a model that costs roughly one-tenth of a frontier model per token while maintaining sufficient quality for routine agent steps, NVIDIA enables a tiered architecture: cheap, fast model calls for the 80% of work that is mechanical, and expensive frontier calls reserved for the 20% that genuinely requires deep reasoning.

The Siemens deployment referenced in NVIDIA’s announcement demonstrates this pattern in industrial settings — running model routing for manufacturing and automation agents where the vast majority of inference steps are predictable and do not need a frontier model’s full capacity.

The Broader Competitive Picture

This release positions NVIDIA beyond its traditional role as a hardware vendor. With Nemotron 3.5 Lightning and NeMo Switchyard, the company is competing directly with OpenAI, Anthropic, and Google not just at the GPU level, but at the model-serving and orchestration layer. The open-source approach is a deliberate contrast: while OpenAI and Anthropic lock customers into proprietary API ecosystems, NVIDIA is giving away the models and the routing infrastructure, betting that openness will drive adoption of its compute platform.

The model is also a response to the open-weight wave from Chinese labs. DeepSeek’s V4-Flash and Alibaba’s Qwen 3.8-Max have dominated the open-weight conversation in recent weeks. NVIDIA — an American company — now offers a credible open alternative specifically optimized for the agent workloads that enterprises are racing to deploy.

The Ecosystem Responds

The day-one availability across multiple inference providers signals strong ecosystem support. OpenRouter listed the model within hours at competitive pricing. DeepInfra offered day-zero access with optimized serving. FriendliAI announced specialized inference optimizations leveraging their serving engine. Amazon SageMaker JumpStart added the model to its catalog. This breadth of deployment options is itself a competitive advantage — enterprises can adopt the model through whichever serving infrastructure they already use.

For developers, the immediate practical takeaway is simple: if you are running autonomous agents and paying frontier-model prices for routine inference steps, Nemotron 3.5 Lightning combined with NeMo Switchyard routing offers a proven path to cutting those costs by more than half — starting today.