Not Just ARC-AGI-3: Quesma's Puzzle Report Shows GPT-6 Astra's Skills Generalize Beyond the Benchmark
A new independent evaluation pits GPT-6 Astra against Claude Fable 5.1 across Portal, Baba Is You and MazeBench — and finds the OpenAI model's puzzle-solving is a general skill, not benchmark overfitting.
When OpenAI’s GPT-6 Astra posted a near-perfect score on ARC-AGI-3 earlier this month, the obvious skeptical question followed within minutes: is this real, generalizable puzzle-solving intelligence, or has the model simply been optimized — “benchmaxxed” — for one specific evaluation? On September 14, 2026, independent evaluator Quesma published the most comprehensive attempt yet to answer that question, and the verdict leans firmly toward “generalizable.”
The report, written by Quesma founder Piotr Migdał, stitches together evidence from wildly different sources: a hobbyist’s 24-hour autonomous playthrough of Valve’s 3D puzzle game Portal, Quesma’s own head-to-head agent evaluation on 154 levels of Baba Is You, a 60-hour crawl through the 3D spatial-reasoning eval MazeBench, and the testimony of Przemysław “Psyho” Dębiak — the puzzle designer who a year ago was the only human to beat OpenAI’s model, and who now struggles to find puzzles Astra cannot solve.
The headline numbers, in context
Start with the benchmark that triggered the debate. On ARC-AGI-3, GPT-6 Astra scores 63% with the standard agent harness — and 99% with a custom one. For comparison, its predecessor GPT-5.6 Sol manages just 8% with default harnesses, and Claude Opus 5 sits at 30%. A 30-plus-point jump over the previous frontier is exactly the kind of leap that makes people suspicious.
But here is what makes Quesma’s report persuasive: the same model, with no harness customization at all, keeps winning at things that are not ARC-AGI-3.
Portal, solved in 23 hours 43 minutes
The most striking datapoint predates the report. In early September, an AI enthusiast known as CozyBlaze let GPT-6 Astra play through the entirety of Valve’s Portal — the 2007 first-person 3D puzzler built on portals, momentum, and spatial reasoning — with no human help. The agent, which controls the game through Model Context Protocol plus a modified SourcePauseTool, had to map 3D spaces from screenshots, understand each chamber’s puzzle, plan a solution, and only then execute its input sequence. The run took 23 hours and 43 minutes, consumed 3,336 tool calls, and would have cost $571.18 in API tokens at list price (CozyBlaze actually ran it inside a $200 Codex Pro subscription).
As Tom’s Hardware, which documented the run, noted: it wasn’t long ago that AI systems were losing at Atari 2600 games. A general-purpose language model autonomously completing a commercial 3D game — one that requires embodied spatial reasoning from 2D screenshots — is a different category of achievement entirely.
Baba Is Astra: 80 levels vs. 33
Quesma’s own contribution to the report is a fresh head-to-head on Baba Is You, the rule-manipulation puzzle game whose levels work like self-modifying logic problems. The team gave agents the 154 levels of the first ten worlds, from The Intro to Volcanic Cavern, using native harnesses — Codex for GPT-6 Astra, Claude Code for Claude Fable 5.1. Each agent had 6 hours, no internet access, and an instruction to solve as many levels as possible.
The result was not close. GPT-6 Astra solved 80 of 154 levels; Claude Fable 5.1 solved 33, completing only two worlds (The Lake and Solitary Island). After three hours, Fable 5.1 declared the nine levels it still had open “infeasible” and stopped on its own — with half its time budget left. Astra went on to solve six of those nine.
The performance details are arguably more interesting than the tally. Astra solved 7 of 70 levels on the first attempt with no undo — including one dispatched with a 79-move sequence just 16 seconds after first seeing the level. And for context on the pace of progress: GPT-5.6 Sol previously burned over 9 hours on The Lake alone and still failed to clear it.
The caveats are stated plainly: Baba Is You is a public game, and parts of it may appear in training data. Quesma checked agent trajectories for mentions of level names or recalled solutions and found none — but as the report acknowledges, the only rigorous way to distinguish memorization from skill is to test on challenges the model hasn’t seen. Which is precisely what the rest of the evidence does.
The human benchmark gives up
That evidence comes from Przemysław “Psyho” Dębiak, the puzzle designer with the unique distinction of having beaten OpenAI’s model a year ago as the only human to do so. He has now tested Astra against around 30 obscure puzzle games — mostly PuzzleScript titles with rule discovery, brutally difficult levels, and minimalist grid-based design, the kind of niche creations that exist nowhere near any training corpus.
His verdict: Astra is “above an average human player, and honestly the ~99% on ARC-AGI-3 might undersell it.”
One pattern across every test
The same shape shows up on MazeBench, a 3D labyrinth of Sokoban-style puzzles that its authors introduced in July 2026 with the admission that “today’s best agents cannot progress beyond the initial levels.” Astra spent over 60 hours in the environment without Python access and finished at 14% — versus 2% for Claude Fable 5.1.
Strung together, the picture is consistent: whether the medium is a benchmark of 2D grid puzzles (ARC-AGI-3), a commercial 3D game navigated through screenshots (Portal), a rule-rewriting logic game (Baba Is You), a 3D open-world planning eval (MazeBench), or an obscure indie puzzle engine (Dębiak’s battery), GPT-6 Astra substantially outperforms every peer — and frequently its own predecessors by an order of magnitude.
Why it matters
This report lands mid-debate for a reason. Benchmark saturation is becoming an existential problem for AI evaluation: the original ARC-AGI-3 was “saturated so surprisingly fast” that, per Dębiak, the ARC team reportedly scrapped the ARC-AGI-4 design and is going straight to ARC-AGI-5. When a benchmark dies in weeks, the only durable signal comes from cross-domain consistency — which is exactly what Quesma assembled here.
The second implication is starker, and the report says it without hedging: “all puzzles (including security locks) that can be solved by experts are easily solved by frontier AI models.” Puzzle-solving was long considered a decent proxy for expert human reasoning. If frontier models now clear that bar across the board, the security community’s assumptions about which locks, challenges, and verification schemes remain “hard for computers” need a rapid refresh.
For Anthropic, the gap is a genuine competitive datapoint: Claude Fable 5.1 remains strong on coding and prose tasks, but on spatial, embodied, and rule-discovery puzzles it is being beaten two-to-one or worse. And for everyone else, Quesma’s parting challenge is the right one: if you know a puzzle game that’s within reach of bright human minds but should be beyond Astra and Fable — that’s now valuable evaluation material. The race to find what frontier models can’t do has become the benchmark game in its own right.
Sources
- [1] https://quesma.com/blog/gpt-6-astra-solves-puzzles/
- [2] https://www.tomshardware.com/tech-industry/artificial-intelligence/openais-gpt-6-astra-model-autonomously-completes-portal-in-24-hours-feat-cost-just-usd571-in-tokens
- [3] https://arcprize.org/blog/astra
- [4] https://openai.com/index/gpt-6-astra/