← All posts / Research

NVIDIA's AVO Agent Aces ARC-AGI-3 With a Perfect Score — By Wrapping Claude Opus 5 in Better Scaffolding

NVIDIA's Agentic Variation Operators architecture scored a perfect 100.00 on the ARC-AGI-3 public set, lifting Claude Opus 5 from a ~30% solo baseline — the strongest evidence yet that agent harness design, not raw model capability, sets the ceiling for long-horizon autonomy.

NVIDIA's AVO Agent Aces ARC-AGI-3 With a Perfect Score — By Wrapping Claude Opus 5 in Better Scaffolding

On August 21, NVIDIA’s research team published a result that reads like a benchmark oddity until you sit with it: an agent system called AVO — Agentic Variation Operators — scored a perfect 100.00 RHAE on the ARC-AGI-3 public benchmark, completing all 183 levels across 25 game environments. The underlying model was Claude Opus 5, the same model that ARC Prize’s own leaderboard puts at roughly 30% when it faces the benchmark alone at High reasoning effort.

No new model was trained. No frontier weights were unveiled. The entire delta between 30% and 100% came from the harness wrapped around the model — persistent memory, a supervisor loop, and an execution architecture designed to sustain progress over long horizons. NVIDIA’s own framing is blunt: evaluating a model is not the same as evaluating an agent.

What ARC-AGI-3 Actually Tests

ARC-AGI-3, the third generation of François Chollet’s abstraction-and-reasoning gauntlet, is an interactive reasoning benchmark. An agent is dropped into an unfamiliar game environment with no instructions, no stated rules, and no stated goal. It must explore through interaction, infer the environment’s dynamics and objectives from the consequences of its own actions, and then plan efficiently across progressively harder levels.

Scoring uses RHAE — Relative Human Action Efficiency — a metric that blends task completion with per-level action efficiency, measured against first-time human baselines. That second part matters: an agent that flails through brute-force exploration eventually finishes levels, but RHAE punishes waste. Success requires preserving useful knowledge across levels, learning from previous interactions, recovering from mistakes, and spending environment actions sparingly.

It is, in other words, a test of exactly the thing current frontier models are worst at: sustaining coherent, compounding progress over a long horizon without a human resetting the context window.

The 30% to 100% Lift

The headline numbers from NVIDIA’s public-set run:

  • 100.00 RHAE across all 25 environments, all 183 levels solved
  • 6,624 environment actions used in total — versus the 7,542 that VISTA, the previous state-of-the-art direct-interaction harness, reports for the same 183 levels with the same Claude Opus 5 backend. That is roughly 12% fewer actions in a cross-system comparison
  • Text-only observations: each game state was rendered as an exact 64×64 text grid, with no images or image tokens sent to the model — a notable simplification versus VISTA’s primary 512×512 rendered PNG configuration

NVIDIA is careful about what these numbers do and don’t prove. The post explicitly notes the comparison to VISTA “should not be interpreted as a controlled ablation” — the two systems differ in agent backend, observation representation, memory, and context management. And the 30% Opus 5 baseline comes from a different reasoning setting and evaluation setup, so the gap illustrates rather than measures AVO’s contribution. The result also covers the public set only, not the semi-private or private competition sets where ARC Prize hands out its prize money.

But even with those caveats, the shape of the finding is hard to argue with: the same model, wrapped in different machinery, produces radically different agent-level intelligence.

Where AVO Came From: Seven Days of GPU Kernels

AVO did not begin life as a game-playing agent. It was built for autonomous GPU-kernel optimization — arguably a harder test of long-horizon discipline, where feedback comes from compilers, profilers, and hardware benchmarks rather than game transitions.

In NVIDIA’s attention-kernel study, AVO ran continuously for seven days, explored more than 500 optimization directions, and committed 40 kernel versions. On DGX B200 systems, the resulting multihead attention kernels beat cuDNN by up to 3.5% and FlashAttention-4 by up to 10.5% across evaluated configurations. The agent then adapted its evolved kernel to grouped-query attention in roughly 30 minutes of additional autonomous work.

The architecture has two load-bearing mechanisms, and they generalize:

Persistent memory carries forward prior implementations, evaluation results, compiler and profiler outputs, and accumulated reasoning — letting the agent resume from its current state instead of reconstructing the search every time the context rolls over. In a benchmark where knowledge from level n is often the key to level n+1, this is the difference between learning and amnesia.

Supervision is a second loop that watches the broader trajectory for stagnation and repeated unproductive cycles, redirecting the main agent toward alternative strategies when progress plateaus. The main agent keeps authority over what to inspect, change, test, and commit; the supervisor’s job is keeping the search moving forward.

NVIDIA’s transfer argument is the interesting part: GPU optimization and ARC-AGI-3 look nothing alike on the surface, but both reduce to the same loop — hypothesize, act through an external interface, observe consequences, preserve useful state, revise the model of the problem, recover from wrong assumptions, and keep going. “What transfers is not domain knowledge,” the team writes, “but the machinery for sustained autonomous progress.”

One Architecture, Multiple Models

AVO is also model-agnostic by design. On a challenging subset of games, NVIDIA paired it with GPT-5.6 Sol, and the preliminary result is a nice illustration of complementary operating profiles: Sol reached matched levels faster in wall-clock time in several cases, while Opus 5 used fewer environment actions in matched-level comparisons. The team leaves a systematic comparison to future work, but the direction is clear — the harness is becoming the portable asset, and models are becoming interchangeable components within it.

Why the Harness-Over-Model Thesis Keeps Winning

AVO is the latest and cleanest data point in a trend that has been building all year. VISTA demonstrated it first with strong results using Claude Code and Codex harnesses; Tycho attacked the same benchmark from the opposite direction with explicit executable world models; and on Prime Intellect’s nanoGPT speedrun published August 23, Claude Fable 5’s win was largely attributed to how well models sustained an autonomous optimize-test-revise loop. The consistent lesson across all of them: a mid-tier model in the right scaffolding outperforms a frontier model running naked.

For developers, this is genuinely good news. It means agent capability is partly an engineering problem — memory design, supervision loops, tool interfaces, recovery strategies — rather than purely a question of who can afford the biggest model. NVIDIA has not said whether AVO will be released as an open framework or remains an internal research demonstration, and that is now the most consequential open question hanging over this result.

There is also a benchmark-governance angle. When harnesses can lift a model from 30% to 100% on a public set, leaderboards that report bare model scores tell you less and less about real-world agent performance. ARC Prize’s semi-private and private sets exist precisely to control for overfitting, and the honest reading of this week’s news is that the public set is now effectively saturated. The frontier has moved: the question is no longer which model is smartest but which system turns model intelligence into sustained, efficient action.

NVIDIA’s conclusion fits on a bumper sticker, and it is one every agent developer should internalize: the model matters, but the model is not the entire agent.