← All posts / Research

8,429 Wrong Pixels to 2: How a Year of Frontier Models Learned to Port Prince of Persia

A developer fed the original 6502 assembly of Prince of Persia to every new frontier model for a year. Claude Opus 5.5 just finished the job — from a single prompt, it ported the game's own room-drawing routine and proved the result pixel by pixel.

8,429 Wrong Pixels to 2: How a Year of Frontier Models Learned to Port Prince of Persia

On September 25, 2026, a blog post by a developer named Priyan quietly hit the front page of Hacker News, and it may be the most legible one-year progress report on frontier coding models anyone has produced. The setup is disarming: take the original Apple II source code of Prince of Persia — the 6502 assembly Jordan Mechner recovered in 2012 from floppy disks that had sat in a box for two decades — and ask each new frontier model to port the game to C#. The human’s only job is to play the result and say what is wrong. No reading the code, no hand-edits. Prompts only.

The punchline is a number. On the first screen of level 1, the count of pixels that differed between the AI’s port and the real game running in DOSBox went from 8,429 to 2 — and the two remaining pixels were a torch flame caught at a different moment of its animation. The model that closed that gap, Claude Opus 5.5, did it from a single prompt. But the story of how four generations of models failed their way to that number is what makes the experiment worth reading.

Round 1 — Opus 4.6: right language, wrong architecture

The first attempt, in February 2026, started with a deliberately modest prompt: “in this original code there 6502 assembly code for prince of persia, use the save level files and try to do it in c# console.” Over two days, Claude Opus 4.6 produced a C# console game, then a level editor, then a Raylib graphics version. It parsed the original level files correctly — rooms, tiles, gates, guard positions — and rendered something that looked like the Apple II game.

It did not play like Prince of Persia. The author’s bug reports from those days read like a slow-motion crash: “graphics is scrambled,” “movement everything wrong … prince is not in floor,” “it dropped below floor and arrow keys doesnt move.” The core problem, diagnosed only later, was architectural. The prince moved one tile at a time, hopping from cell to cell. The real game plays hand-drawn rotoscoped animation frames with a small movement per frame. Opus 4.6 could read 6502 and write C#, but it picked the wrong engine and could not tell.

Round 2 — OpenAI Codex: polish on a broken foundation

In March 2026 the same codebase went to OpenAI’s Codex: “this is prince of persia original assembly code to c# port done by claude still graphics not good can you fix it.” Codex made genuinely sensible small fixes — sharp pixel filtering so the art was not blurred, transparent sprite backgrounds, open gates you could walk through, pillar tops that stopped acting like walls. The prince still ended up in wrong positions and still got stuck. It polished the surface without ever questioning the foundation, and it never ran the game.

Round 3 — Opus 5: the model that could see

September 2026 brought a rule change that mattered more than any model upgrade. The author wrote two Claude Code skills — one to control DOSBox (launch programs, send keys, take screenshots) and one to control Windows programs — then pointed Opus 5 at his original DOS copy of the game with an instruction to compare against the real thing and keep going until it succeeded or 6am.

It worked through the night, and the first thing it did was diagnose the tile-grid architecture problem and rebuild the engine around the original’s frame-sequence design. Then it went further: it reverse-engineered the DOS game’s file formats and read the real data straight out of them — every animation frame of the prince, the dungeon art, all 15 levels. It located the original animation tables inside PRINCE.EXE by searching for byte patterns recognizable from the Apple II source, and confirmed each table against a second known value before trusting it. By morning the game was playable with genuine animation. In a later session the author played the first screens in DOSBox while the model watched the screenshots and found its own ledge-detection bug — then fixed it by going back to the original assembly routine.

The levels still looked wrong. Bricks, gates and decorations were slightly off, because Opus 5 had placed every piece by matching screenshots one at a time, filling whatever it could not explain with guesswork.

Round 4 — Opus 5.5: find how the original actually does it

When Opus 5.5 shipped, it got a single prompt: “the character moves runs jumps but gate positions bricks all different, you are new model let see anything u can improve.”

It took a categorically different approach. Instead of tuning positions by eye, it went looking for how the original game actually draws a room — and found it documented in SDLPoP, the open-source port of the DOS game that Dávid Nagy and the princed.org community reconstructed from its disassembly over years of work. Opus 5.5 ported SDLPoP’s room-drawing routine (src/seg008.c) line for line into a C# file, DosRoomDrawer.cs, whose header openly credits the source and the GPL-3.0 license.

That single move explained things the earlier sessions had only guessed at. The “random” brick pattern is not random: the original seeds its RNG from the room, row and column, so every wall looks identical on every visit. The gate at the start of level 1 is actually open in the level data — the game presses a hidden button as you drop in, which is why you hear it slam shut behind you.

Opus 5.5 also caught two of its predecessor’s errors on its own: PRINCE.EXE is compressed with Microsoft’s EXEPACK, meaning the tables Opus 5 found had only worked because they happened to sit in a stretch the packer left untouched — so it wrote an unpacker and read the tile-drawing tables properly. And an earlier theory about “stencil colours” was wrong; those colours are the real teal of the exit door.

Then it verified itself, checking pixel by pixel against the real game in DOSBox: 8,429 differing pixels down to 2, with the second room and the level 3 exit door matching as well. It was not flawless — a sprite-position change it had verified only for one facing direction made the prince sink into walls when flipped, and the author restored a backup — but even the failure mode was new: the model proved the movement itself was identical across every test tick, isolated the sprite bug, and reverted exactly that part.

The honest caveat, in the model’s own words

To his credit, the author asked the obvious question: could Opus 5.5 have worked this out from the DOS binary alone? The answer he got is worth quoting, because models that assess their own capabilities this precisely are new:

“Honestly, probably not in one prompt. I didn’t reverse-engineer the drawing code. I read SDLPoP’s reconstruction and ported it. I only found the tables inside PRINCE.EXE because SDLPoP told me what their contents were. The previous model searched the executable without that key and found nothing. … What I did bring was knowing that SDLPoP existed, reading it correctly, spotting that the EXE was compressed, and proving the result pixel by pixel.”

Why this matters beyond retro games

Strip away the nostalgia and the experiment isolates three variables that the benchmarks miss. First, architecture beats polish: Codex’s pixel-perfect filtering on top of a tile-grid engine was worth nothing, while finding the frame-sequence design was worth everything. Second, verification is the real unlock — from Opus 5 onward, the model could see the original and test itself against it, and the human stopped being the only tester. Third, knowing where the knowledge lives is becoming a first-class capability: Opus 5.5’s edge was not raw reasoning over a binary but recognizing that SDLPoP’s decade of reverse-engineering existed, reading it correctly, and porting it with attribution.

There is also a cautionary thread for anyone tracking “AI can write software” claims. The winning move leaned on a GPL-3.0 licensed community reconstruction — and the model handled the licensing honestly, carrying the notice into the ported file’s header, which forced the whole project to GPL-3.0-or-later. Provenance and license hygiene are now model behaviors, not just human ones.

The experiment is not over. Guards and sword fighting, the palace levels, and exact landing positions remain. The author says he will run the same experiment on the next frontier model and see how far it gets. The full session history is on GitHub under GPL-3.0, which means the next model — and the next — will be able to read everything their predecessors did wrong.