NVIDIA's Nemotron 3.5 Lightning Is a 30B Open Model Built to Do the Agent Grunt Work
NVIDIA's new open 30B MoE model with just 3B active parameters targets the high-volume execution layer of always-on agents — 4x faster output, 54.3% on SWE-Bench Verified, and a 1M-token context, all under the permissive OpenMDW license.
Here is the uncomfortable truth about agentic AI in 2026: most of the tokens your agents burn are not spent on brilliant reasoning. They are spent on the repetitive middle — summarizing a pull request, classifying an alert, validating a tool output, formatting a result, delegating to a subagent. Routing every one of those steps to a frontier reasoning model, as many teams still do, is like dispatching a container ship to deliver a single package.
On August 11, NVIDIA shipped its answer: Nemotron 3.5 Lightning, an open 30B mixture-of-experts model with just 3B active parameters, purpose-built for the execution layer of always-on agents. It is the smallest member of the Nemotron 3 family, distilled from the 550B-parameter frontier model Nemotron 3 Ultra, and it arrives alongside NeMo Switchyard, an open-source model-routing library that lets agent systems send each request to the right model for the job.
What it is
Nemotron 3.5 Lightning is a hybrid MoE: 30B total parameters, but a router sends each token to only a few experts, so roughly 3B parameters actually compute per token. That is the core trick behind the model’s economics — the capacity of a larger dense model at something close to the compute cost of a small one.
The headline specs:
- Architecture: hybrid Mixture-of-Experts, distilled from Nemotron 3 Ultra, developed with the “Nemotron Coalition” of partners
- Active parameters: ~3B per token out of 30B total
- Context length: up to 1M tokens — enough to hold a full repository, an alert history, or an entire multi-turn session state without truncation
- Precision: ships with both BF16 and NVFP4 checkpoints, using the same specialized NVFP4 kernels that power Nemotron 3 Ultra across Blackwell, Hopper, and even Ampere GPUs
- License: OpenMDW-1.1 — weights, training data, and recipes released as permissively as NVIDIA can make them
That last point deserves emphasis. In a year when “open” has often meant open weights and nothing else, Nemotron 3.5 Lightning ships the full stack: LoRA and full-SFT recipes via NeMo Automodel and NeMo Megatron Bridge, reinforcement-learning tooling via NeMo RL and NeMo Gym, and even an open agentic RL dataset — Nemotron-RL Agentic Terminal Pivot — used to train some of the model’s coding-agent capabilities. Small models also fine-tune faster, cheaper, and on far more modest hardware, which is precisely what makes this tier of model attractive to customize.
Speed as a feature
Speed is the headline claim. NVIDIA says Lightning delivers up to 4x faster output compared to similar-sized models, and completes agentic tasks up to 30% faster. On PinchBench, the model reaches 86% accuracy while completing 10,000 tasks 30% faster than Qwen3.6 35B at similar accuracy. On the Artificial Analysis Intelligence Index — a combination of nine evaluations spanning agentic tasks, coding, scientific reasoning, and general intelligence — NVIDIA claims the model defines the accuracy-versus-speed Pareto frontier for small open models.
The speed comes from three techniques working together. First, multi-token prediction (MTP) was baked into pretraining, with a dedicated MTP-boosting phase afterward, letting the model draft multiple tokens per forward pass. Second, speculative decoding: the release ships with two draft models — DSpark, recommended for DGX Spark and low-concurrency data center workloads, and DFlash, which teams can benchmark against the others for their own workloads. Third, the NVFP4 quantized checkpoint shrinks memory traffic on modern GPUs without a separate conversion step.
The generational accuracy leap
The comparison against Lightning’s predecessor, Nemotron 3 Nano, is stark. On SWE-Bench Verified — the benchmark that measures whether a model can actually resolve real GitHub issues — Lightning scores 54.3% versus 21% for the previous generation. On GDPval-AA v2, a knowledge-work benchmark, it rates 924 ELO versus 479. That is more than double the coding score and nearly double the knowledge-work rating, alongside what partners report as the strongest hallucination-resistance numbers in its comparison set. NVIDIA notes these are preliminary results from a model still in training, with final GA numbers expected to improve.
NVIDIA attributes part of this to harness-optimized training: the model is explicitly trained for popular agent harnesses — the blog names OpenClaw and Hermes Agent, supported by the NemoClaw open-source security and management stack — so it makes more accurate tool calls while keeping latency low on high-volume tasks.
One system of models, routed by Switchyard
The deeper story is architectural. NVIDIA is explicit that developers increasingly build with a system of models: frontier reasoners like Nemotron 3 Ultra handle orchestration and complex planning, while smaller, faster models handle the execution layer. NeMo Switchyard, released alongside Lightning, makes that division of labor practical. It exposes Lightning as a routing target next to your open and closed models, so plans route up to the frontier and execution routes down to Lightning. Because both ends speak the same OpenAI-compatible API, a router can move traffic between them without changing application code.
The partner ecosystem around the launch signals how seriously NVIDIA is pushing this: harnesses and agent frameworks including Cline, Factory AI, Kilo Code, LangChain, OpenClaw, OpenCode, and OpenHands; inference software from Ollama, LM Studio, llama.cpp, and Unsloth; hosted inference from Baseten, CoreWeave, DeepInfra, Fireworks AI, GMI Cloud, Modal, Nebius, and Together AI; and cloud platforms spanning SageMaker JumpStart, Google Cloud, Microsoft Foundry, and OCI.
Local-first agentic AI
Lightning is also a local-AI play. The model runs on NVIDIA Jetson, GeForce RTX 5090, and DGX Spark — capable agentic AI on a desktop, in other words, not just a data center. On the EXO Labs local.ai leaderboard, Lightning again sits on the Pareto frontier for small open models. For teams worried about sending agent traffic to external APIs, a 30B MoE with 3B active parameters is exactly the size class that makes on-premise, always-on agents economically viable.
Why it matters
The frontier-model discourse fixates on leaderboard-topping reasoners, but the economics of deployed agents are decided in the trenches — in the millions of cheap, fast, reliable calls that dominate a long-running agent’s token budget. DeepSeek’s weekend move to peak-hour billing only underscores how much the cost of frontier tokens is now in flux. A fully open model with a 1M-token context, frontier-class distillation, honest benchmark gains, and a permissive license changes the build-vs-buy calculus for anyone assembling agent fleets in 2026.
Nemotron 3.5 Lightning is available now on build.nvidia.com and OpenRouter, with weights on Hugging Face and ModelScope. If your architecture still routes every tool call to a frontier model, this release is NVIDIA’s politely worded suggestion that it might be time to stop.
Sources
- [1] https://developer.nvidia.com/blog/nvidia-nemotron-3-5-lightning-delivers-fast-accurate-specialized-task-execution-for-long-running-agents/
- [2] https://blogs.nvidia.com/blog/nemotron-lightning-switchyard-rtx-dgx/
- [3] https://www.gmicloud.ai/en/blog/nvidia-nemotron-3-5-lightning-is-live-on-gmi-cloud-what-your-agent-system-was-missing
- [4] https://openrouter.ai/nvidia/nemotron-3.5-lightning