30 Out of 42, Fully Open-Sourced: NVIDIA Publishes the Complete Recipe Behind Nemotron's IMO Gold
NVIDIA releases the checkpoints, training data, inference code, submitted solutions and a 200-problem benchmark behind its natural-language pipeline that scored IMO-gold 30/42 with no formal prover.
The race to gold-level mathematical olympiad performance has produced plenty of impressive scores this year, but almost every one of them came wrapped in a black box. NVIDIA just did the opposite. In a paper titled “An Open Recipe for IMO Gold: Training Nemotron for Olympiad Mathematics” (arXiv:2609.10712), a team led by Ivan Moshkov and Igor Gitman has published not just the score — 30 out of 42 points at the International Mathematical Olympiad 2026, above the 29-point gold threshold — but the entire machinery that produced it: two post-trained checkpoints, the training data, the training and inference code, the actual solutions submitted during the competition, and Nemotron-IMO-Bench, a brand-new benchmark of 200 novel olympiad-level problems.
The release landed quietly on Hugging Face this week as the “Nemotron Labs IMO 2026” collection, and it deserves attention precisely because of what it refuses to hide.
What the system actually is
Strip away the branding and the pipeline is conceptually simple, which is part of why it matters. The system starts from Nemotron 3 Ultra — the same open-weights family NVIDIA has been shipping all year — and trains two specialist checkpoints on top of it, one via supervised fine-tuning on curated proof data, and one via reinforcement learning on top of that. The paper’s core question is refreshingly empirical: how much of olympiad-level performance comes from post-training, and how much comes from what you do at inference time?
The answer turns out to be: both, inextricably. Three Nemotron 3 Ultra checkpoints — the general-availability model plus the two specialists — power an iterative search loop that generates candidate proofs, attempts to verify them, and refines the failures. A separate high-compute selection stage then picks the final submission for each problem from everything the search produced.
Two constraints make the result notable. First, the system operates entirely in natural language: no Lean, no Isabelle, no formal theorem prover of any kind. Second, no external tools and no internet access — the model cannot phone a friend, call a symbolic algebra engine, or look anything up. Every proof is generated, checked, and polished by the models themselves, in plain mathematical English.
Why natural-language-only matters
Most headline olympiad results from frontier labs have leaned on formal methods or tool-augmented pipelines, where a verifier can mechanically confirm each step. That approach buys rigor but narrows what the system can attempt: formalization is expensive, and many competition problems resist clean encoding. A natural-language pipeline has no such guardrail, which makes both the gold score and the published verification-and-refinement loop more interesting — the team explicitly studied checkpoint choice, verification, and refinement as independent variables and reports what each contributed.
For the open-source community the implications are direct. Anyone with sufficient compute can now reproduce the approach, ablate it, or push it further. The two specialist checkpoints are downloadable. The RL dataset (Nemotron-Math-Proofs-v3-RL and its predecessors in the collection) is downloadable. The inference code that orchestrates generate-verify-refine is downloadable. And crucially, the submitted solutions from the actual IMO 2026 run are in the drop, so the score itself can be audited rather than taken on faith.
Nemotron-IMO-Bench: 200 problems that aren’t anywhere else
The quiet star of the release may be the benchmark. Contamination is the chronic disease of math evaluation: any problem that has circulated online is already in every frontier model’s training data, which makes reported scores progressively less meaningful. Epoch AI’s recent declaration that FrontierMath Tier 4 is saturated — GPT-6 Astra moved the top score from 5% to 98% in fourteen months — is the latest symptom.
NVIDIA’s answer is Nemotron-IMO-Bench: 200 novel, previously unpublished olympiad-level problems written for this release. A fresh benchmark of that difficulty, released alongside a reproducible gold-level pipeline, gives the community both a measuring stick that hasn’t leaked into training corpora and a known-good baseline to measure against. Expect every serious math-reasoning effort to report on it within months.
Context: open weights are closing the gap
The release also lands amid a broader shift. In the same week, Epoch noted that GPT-6 Astra’s FrontierMath saturation involved models “exploiting unintended shortcuts” on several problems — exactly the kind of subtle failure that only transparent pipelines and uncontaminated benchmarks can surface. And the open-weights side of the aisle keeps compounding: Moonshot’s Kimi K3 at 2.8T parameters is now the substrate for Cognition’s near-frontier SWE-2 coding agent, and NVIDIA itself has spent 2026 shipping Nemotron variants for agents, code, and reasoning on a near-monthly cadence.
Against that backdrop, publishing a complete IMO-gold recipe reads less like a flex and more like a strategic bet: the moat in olympiad math is evaporating, and the value is migrating to whoever can prove — reproducibly — what their systems actually do. NVIDIA’s bet is that being the lab that shows its work wins the long game.
What to do with it
For researchers, the paper is a clean ablation study of test-time compute in a domain where verification is genuinely hard. For engineers, the collection is a working template for building iterative generate-verify-refine systems on open models — a pattern that transfers well beyond math to code review, formal verification of designs, and any domain where a model must check its own work. For everyone else, it is a reminder that the frontier of mathematical reasoning is no longer sealed inside a handful of labs.
The full collection lives at huggingface.co/collections/nvidia/nemotron-labs-imo-2026, the paper at arXiv:2609.10712. Thirty out of forty-two points, no prover, no tools, no secrets.