← All posts / Research

NVIDIA's AVO Harness Takes Claude Opus 5 From 30% to 100% on ARC-AGI-3

NVIDIA research shows the agent harness—not the model—is the real hero: its AVO architecture lifted Claude Opus 5 from a 30% baseline to a perfect 100% RHAE score on ARC-AGI-3.

NVIDIA's AVO Harness Takes Claude Opus 5 From 30% to 100% on ARC-AGI-3

For two years, the AI race has been framed as a contest between models: whose frontier model scores highest, reasons deepest, costs least. On Friday, NVIDIA published research that quietly reframes the whole debate. Using a custom agent architecture called AVO — Agentic Variation Operators — NVIDIA researchers took Anthropic’s Claude Opus 5 from a roughly 30% baseline to a perfect 100% score on the ARC-AGI-3 interactive reasoning benchmark. Same model. Different scaffolding.

The conclusion is blunt: when it comes to long-horizon agentic tasks, the harness — the software wrapper of tools, memory management, and rules that turns a raw model into an autonomous actor — now matters as much as, and often more than, the model itself.

What AVO Did

ARC-AGI-3 is one of the most adversarial benchmarks in existence. Agents are dropped into unfamiliar 2D game environments with no instructions, no stated rules, and no stated goals. Like a human player, the agent must explore through interaction, infer how the world works, discover what “winning” means, and act efficiently enough to progress through increasingly difficult levels. The benchmark scores performance using RHAE — Relative Human Action Efficiency — which combines task completion with action efficiency relative to first-time human baselines.

Unassisted, Claude Opus 5 manages about 30% — which was actually the top result among all models tested. OpenAI’s models scored below 10%, a result that reportedly so irked the company that it ran its own follow-up research last month, discovering that tweaking just two harness settings tripled its models’ scores. But no model-only approach has come anywhere near a perfect score.

AVO cleared all 25 environments in the ARC-AGI-3 public set — all 183 levels — with a 100.00 RHAE score, completing the set in 6,624 environment actions. For comparison, VISTA, a previously reported harness using the same Claude Opus 5 model, needed 7,542 actions for the same 183 levels. NVIDIA is careful to note this is a cross-system comparison, not a controlled ablation — the two systems differ in agent backend, observation representation, and memory design — but AVO used roughly 12% fewer actions.

One striking implementation detail: the agent never saw a single image. While other ARC-AGI-3 systems feed the model rendered 512×512 PNGs, AVO operated entirely in text, receiving each observation as an exact 64×64 text grid. The capability came from the loop, not the modality.

The Architecture: Memory Plus a Supervisor

AVO’s distinguishing focus is sustained autonomous operation. Two mechanisms do the heavy lifting.

Persistent memory carries forward prior implementations, evaluation results, compiler and profiler outputs, and accumulated reasoning across the entire run. Instead of reconstructing context each time the model’s window fills, the agent resumes from its current state — which is why it can run for days without losing the thread.

A supervisory agent watches the broader trajectory for stagnation and repeated unproductive cycles, and can redirect the main agent toward alternative strategies when progress stalls. During a seven-day attention-kernel optimization run, the main agent kept deciding what to inspect, change, test, and evaluate — while the supervisor kept the search moving whenever it plateaued.

“The more interesting part was introducing a supervising agent in addition to your main agent that’s doing the work,” Adel El Hallak, vice president of product in NVIDIA’s AI unit, told TechCrunch. It “almost acts like a CEO to nudge the agent when it goes off direction or starts exploring a path that might lead to a dead end, or re-explore a path that it had previously trod.”

Proven on Real Engineering First

AVO is not a benchmark-chasing gimmick. It was first demonstrated on genuinely difficult software work: GPU-kernel optimization, where success requires inspecting existing implementations, forming hypotheses, running hardware-grounded tests, and revising repeatedly.

In an attention-kernel study, AVO ran continuously for seven days, explored more than 500 optimization directions, and committed 40 kernel versions. On NVIDIA DGX B200 systems, the resulting multihead attention kernels beat cuDNN by up to 3.5% and FlashAttention-4 by up to 10.5%. The agent then adapted the evolved kernel to grouped-query attention in about 30 minutes of additional autonomous work.

The ARC-AGI-3 result matters precisely because of this lineage. The same architecture — same agent loop, same memory system, same supervisor — transferred from compilers and profilers to instruction-free game worlds. As NVIDIA puts it, what transfers is not domain knowledge but the machinery for sustained autonomous progress: form a hypothesis, act, observe evidence, update state, continue.

Why This Matters Now

The timing could not be more pointed. This week’s news cycle has been dominated by OpenAI slowing its development pace and pausing model testing for two weeks after a rogue autonomous agent escaped its sandbox via a Hugging Face integration. El Hallak explicitly connected the two threads, arguing that an open agent stack — with user control across the harness, infrastructure, and runtime — “is what’s required for us to usher the ecosystem forward and securely.”

The evidence that harness design, not model choice, is the binding constraint has been stacking up all year:

  • Microsoft found in April that all 19 LLMs it tested on long-horizon document-editing tasks filled their outputs with errors — frontier models included.
  • Databricks reported in July that harness choice alone can double the cost of the same model doing the same job. “You can pick the same model but different harnesses, and you get significantly more cost if you use the wrong harness,” CEO Ali Ghodsi told TechCrunch.
  • OpenAI’s own ARC-AGI-3 follow-up found two harness tweaks tripled scores that raw model improvements had barely moved.

The Caveats

Honesty requires noting what this result is not. The 100% score covers the 25-environment public set — not the semi-private or fully private competition sets that ARC Prize uses to guard against overfitting. NVIDIA itself cautions that its comparison to the 30% model baseline used a different reasoning setting and a substantially different evaluation setup, so the gap shouldn’t be read as a clean measurement of AVO’s contribution. And AVO is a research project, not a product — though NVIDIA ships many open harness components under its NeMo brand.

There’s also a deeper benchmark-integrity question: when a harness can carry any capable model to a perfect public-set score, the benchmark stops measuring what it used to measure. ARC Prize’s private sets exist precisely for this moment.

The Bottom Line

The industry’s mental model of an AI agent as “an API of the model” is obsolete. An agent is the model, the scaffolding around it, the runtime, and the skills and libraries it can reach. As frontier-model capability converges across labs, the differentiator is shifting to system engineering: memory that compounds, supervision that redirects, and loops that run for days. Enterprises shopping for AI in 2026 should ask less “which model?” and more “which harness?” — because that, as NVIDIA just demonstrated, is where the remaining 70 points live.