535 Out of 600: NVIDIA's Nemotron Becomes the First AI to Beat Every Human at IOI 2026
NVIDIA's open-weight Nemotron-3-Ultra-CC scored 535.4/600 at IOI 2026 — beating the top human (498.27) and gold threshold (361.12) live, under identical contest constraints, using GenCorrect feedback-driven test-time refinement.
For nearly four decades, the International Olympiad in Informatics (IOI) has been the proving ground for the world’s most talented young programmers — high-school students who spend years mastering algorithmic reasoning that most professional engineers never attempt. On September 2, 2026, that era quietly ended. A paper published on arXiv by NVIDIA researchers documents what the authors believe is the first AI system to outscore the highest-scoring human contestant on an IOI problem set: Nemotron-3-Ultra-CC scored 535.4 out of 600 at IOI 2026, surpassing the top human score of 498.27 by 37.1 points and clearing the gold-medal threshold of 361.12 by more than 174 points.
What makes this result different from the usual stream of “AI beats humans at X” headlines is the conditions under which it was achieved. This wasn’t a retrospective benchmark run on already-public problems. The system ran live during the actual competition, before the problems were publicly released, under the same constraints as human contestants: no internet access, local code execution only, and a hard limit of 50 submissions per problem with at most one submission per minute. It was a strictly prospective evaluation — no retries, no post-hoc tuning, no second chances.
The pipeline: 22,000 problems, 1.2 million synthetic traces
The paper, “Post-Training Language Models for Gold-Medal Performance in Coding Competitions” (arXiv:2609.02849), is authored by Aleksander Ficek, Sean Narenthiran, Mehrzad Samadi, Somshubra Majumdar, and Boris Ginsburg. Its stated goal is methodological transparency: previous AI competitive-programming results — OpenAI’s ICPC experiments, DeepSeek’s, and others — were often closed systems that bundled together changes in data, scale, post-training, and inference-time compute, making it impossible to isolate which ingredients actually mattered.
NVIDIA’s pipeline is fully dissected:
- Problem curation. 22,000 curated competitive-programming problems, each packaged with test cases, auxiliary files, and reference solutions. Only environments producing consistent verdicts across reference and generated solutions were retained. All IOI 2025, ICPC 2025, and LiveCodeBench Pro problems were excluded from training and deduplicated against the evaluation sets.
- Synthetic reasoning traces. DeepSeek-V4-Flash generated 1.2 million reasoning traces to train the compact model and 477,642 traces for the larger one.
- SFT + RL. Nemotron-3-Nano-CC (30B total, 3B active parameters) received both supervised fine-tuning and GRPO-based reinforcement learning. Nemotron-3-Ultra-CC (550B total, 55B active) received SFT alone — RL at that scale exceeded the available compute budget.
The RL setup is worth noting for its simplicity: 16 rollouts per prompt across 64 prompts (1,024 rollouts per step) at temperature 1.0, with generated C++17 solutions compiled and executed for a terminal reward of 1 for full credit and 0 otherwise. No reference-policy KL penalty, token-level clipped policy gradient — lean, verifiable-reward RL.
GenCorrect: the test-time multiplier
The single most interesting contribution is GenCorrect, a feedback-driven test-time compute strategy that iteratively generates, evaluates, and refines diverse solutions. Rather than sampling once and hoping for the best, the system runs multiple rounds that use feedback from its own prior submissions to steer the next batch of candidates — generating large pools (200 solutions per problem in early rounds, 1,000 in the final round), clustering them for diversity, and using score-blind ranking to pick what to submit within the 10-submission-per-problem practical budget.
The ablation numbers show how much each stage contributes. On IOI 2025, the base Nemotron-3-Nano scored 130 points. After SFT: 280. After RL: 291. After five rounds of GenCorrect: 468 — comfortably above the IOI 2025 gold threshold of 438.3, using a model with only 3B active parameters. The general Nemotron-3-Ultra-CC reached 502 on the same problems. The pattern in the authors’ own conclusion: SFT provides the largest single-sample gains, RL adds a smaller increment, and GenCorrect multiplies what the trained model can do.
For the final round’s expanded selection, the team borrowed inspiration from GenCluster: the model itself was prompted to produce 50 problem-specific test-input generators and validators, executed and filtered until 100 valid test inputs remained, then used for execution-based ranking of the 1,000-candidate pool. The system effectively built its own private test suite for each problem before choosing what to submit.
Engineering under contest rules
Live deployment demanded serious systems engineering. During IOI 2026, the system ran on a peak allocation of up to 760 NVIDIA GB300 GPUs. Quantization trade-offs were measured explicitly: the live configuration used NVFP4 precision with an FP8 KV cache, prefix caching disabled, and multi-token prediction depth of 5 — achieving 52.8% Score@1 at 736.8 tokens/s/GPU. Compared with the BF16 baseline, that sacrifices 6.6 percentage points of single-sample accuracy for a 3.7× throughput increase, which is what made the large candidate batches GenCorrect needs feasible inside a five-hour contest window.
Two further competition-specific adaptations: SFT for the Ultra model used GLM-5.2-generated data (GLM-5.2 beat DeepSeek-V4-Flash as an SFT teacher on IOI 2025), and the team constrained output length, since shorter outputs permit more candidates within the fixed inference window.
Why it matters
Three implications stand out.
1. The “open vs. closed” framing keeps eroding. Nemotron-3 is an open-weight family, and this paper releases the full recipe — data curation, SFT, RL, and test-time strategy — under CC BY 4.0. The first system to beat every human at IOI is one anyone can download, inspect, and reproduce.
2. Test-time compute is now inseparable from “the model.” The same weights scored 304 (Ultra base), 502 (with five GenCorrect rounds), and 535.4 (competition-tuned live run). When we say “model capability” in 2026, we increasingly mean a system property — weights plus inference strategy — not a property of the checkpoint alone.
3. Prospective evaluation is the gold standard now. Because the system ran before the problems were public, contamination concerns are structurally eliminated. That’s the methodology regulators, enterprise buyers, and benchmark designers should demand more of.
The caveats
Honest accounting requires noting the asymmetries that remain. The human contestants are high-school students working alone on a single machine; Ultra-CC is a 550B-parameter model backed by up to 760 datacenter GPUs and a five-figure submission strategy optimized over a year of development on IOI 2025. “Under the same contest constraints” is true of rules, not of resources. The authors themselves frame the contribution as isolating which components matter, not as claiming human parity in any general sense.
And IOI is a specific kind of test — algorithmic puzzles with verifiable answers and clean reward signals. It rewards exactly the kind of verifiable-reward RL and execution-based refinement this pipeline uses. Whether the same stack translates to ambiguous, real-world software engineering remains a genuinely open question.
Still, the milestone is real, and the paper behind it is unusually readable and complete. The era of AIs merely medaling at olympiads is over; the first one has now topped the scoreboard.