Fewer Moves Than a Human: ARC Prize's Independent GPT-6 Astra Analysis and the AGI Forecast It Moved
ARC Prize's neutral re-run of GPT-6 Astra scores 62.7% on ARC-AGI-3 — but beats the median human in action count on 96% of levels, invents its own algebraic notation, and flips the thinking-vs-cost curve, pulling François Chollet's AGI timeline forward.
When OpenAI launched GPT-6 Astra on September 3, the headline number was 99.9% on ARC-AGI-3 — a near-perfect score on the interactive reasoning benchmark that François Chollet’s ARC Prize built to measure the “residual gap” between AI and general intelligence. But that number came from a harness OpenAI designed itself. On September 3, ARC Prize published its own independent analysis, and the two numbers that matter most from it are not the headline score at all.
The first: under ARC Prize’s provider-neutral Standard harness, Astra scores 62.7% for roughly $26,000 — state of the art, but a long way from 99.9%. The second, and the genuinely historic one: in the runs with OpenAI’s harness, Astra used fewer actions than the median human player on 96.0% of levels, averaging 51.7% fewer actions per level. By the benchmark’s measure of action efficiency — how much experience with an environment a solver needs before it masters it — a frontier model has matched and surpassed human parity for the first time.
Two Harnesses, Two Very Different Scores
The gap between 62.7% and 99.9% is the story of two evaluation philosophies. ARC Prize’s Standard harness gives every model the same minimal interface: all the information needed to solve each game, but the model alone decides what to preserve in its visible notes between turns. It is the apples-to-apples comparison the leaderboard is built on.
The Provider Adapter harness, by contrast, lets Astra use OpenAI’s own context-management machinery — preserving opaque reasoning state between requests and compacting long conversations. Under those conditions Astra (high) hits 99.9% for $19K, and the runs were roughly 3.66x faster by elapsed time and used 49% fewer tokens across the 167 game-reasoning pairs both harnesses solved.
ARC Prize is transparent about the tension: only the Standard-harness figure allows a fair cross-vendor comparison today, though the organization now plans to publish vendor-harness numbers on the leaderboard too, clearly labeled as a separate evaluation condition. For context on how far the field has moved: GPT-5.6 Sol managed 7.78% on the same setup, and Anthropic’s Claude Opus 5 sits at 30.16%.
The Result Nobody Predicted
Before launching ARC-AGI-3, ARC Prize tested roughly 500 members of the general public to establish a human baseline. For each level, the baseline is the median action count among players who solved it. The organizers’ working hypothesis was that action efficiency would remain a durable dividing line between humans and AI — that even when a model solved an environment, it would need substantially more flailing exploration than a person.
That hypothesis is now dead, at least for frontier models. The pattern ARC Prize observed is almost binary: once Astra “understands” a game’s mechanics, its execution lands within — and usually below — the human efficiency range. Brute-force approaches still waste actions, but Astra doesn’t brute-force. It builds a model of the world, then acts on it.
Thinking Harder Makes It Cheaper
One of the strangest findings in the data is a cost inversion. Normally, cranking up reasoning effort makes a benchmark run more expensive — more thinking, more tokens. With Astra on the Standard harness, the opposite happens: costs fall from $49,791 with no reasoning to $26,098 at maximum effort, while the score climbs from 35.2% to 62.7%.
The explanation is mechanical once you see it: Astra solves games in fewer moves at higher effort, and fewer moves mean fewer model calls and fewer tokens. The compute moves from exploration to deliberation, and deliberation turns out to be the cheaper resource. (One oddity: the “low” setting scores 17.5%, worse than no reasoning at all — possibly because Astra’s new architecture can loop processing internally before its first token, making a little bit of explicit reasoning the worst of both worlds.)
A Model That Writes Its Own Algebra
The replays are where Astra’s behavior becomes striking. On the Standard harness, everything not saved to visible notes is lost between turns — so Astra developed an on-the-fly, domain-specific shorthand that Chollet describes as “essentially a game-specific algebraic notation.” Examples from the logs: L8: hub q2 (8↓). Lengths: 14=1… records the level, a rotation index, and mechanism lengths; extend8 to3; retract10 to2; shorten8 to1 is an ordered multi-step plan; Turn 5: P=(24,20), empty, facing west carries full state and orientation.
Chollet’s read on X is the most quotable line in the report: Astra exhibits “symbolic modeling behaviors we had previously only seen with sophisticated harnesses, so harness capabilities are increasingly shifting into the model itself.” The scaffolding is migrating into the weights.
In the third-party PRO-LONG harness, which adds a code sandbox, Astra went further — writing small per-game software libraries. In one maze game with guards and patrols it produced maze_solver.py, then combat_solver.py, then patrol_solver.py, plus a sync_state.py that continuously checked its predictions against observations. (ARC Prize notes these runs measure model-plus-tools, not the model alone, since human testers had no interpreter. It also observed no sandbox-escape attempts.)
Chollet Moves His Forecast
ARC Prize is emphatic that saturating ARC-AGI-3 is not proof of AGI: the environments are deterministic, closed-ended, and orders of magnitude shorter than real-world tasks. “All we know about the system so far are its benchmark scores,” Chollet writes.
But the timeline has slipped. When ARC-AGI-3 launched about six months ago, Chollet estimated “about a year” to saturation. Astra arrived in roughly half that. Asked whether his 2030 AGI forecast still held, he replied: “Sooner, because progress is happening faster than I expected.” ARC-AGI-4, in development since ARC-AGI-3’s release and targeted for Q1 2027, will push into recursive self-improvement and open-ended innovation.
Context: The Benchmark Picture Is Messy
The independent analysis lands amid contradictory verdicts elsewhere. Epoch AI’s composite index (50+ benchmarks) puts Astra clearly first with 169 points; Artificial Analysis rates it at 61 — level with its predecessor Sol and behind Claude Fable 5.1 at 66. Astra costs 2.5x more per token than Sol, yet needs about a third of Sol’s compute steps per task. Its AA-Omniscience hallucination rate dropped from 92 to 51 percent, it lost roughly 80 Elo on GDPval-AA v2, and on FrontierMath Erdős it was the only model to solve two of 68 open Erdős problems with Lean-verified proofs at $300 per attempt.
The through-line from ARC Prize’s data, though, is hard to argue with: the thing that made harnesses valuable — persistent state, compact notation, deliberate planning — is becoming a property of the model itself. That, more than any single score, is what moved an AGI forecaster’s pencil.