← All posts / Research

Clone the App, Not the Code: Microsoft's ProgramDistill Turns Working Software Into 4,063 Verifiable Coding Tasks

Microsoft Research's mine-craft-patch pipeline extracts 1,975 replay-verified behaviors from 26 working web apps to auto-build 4,063 SWE tasks with zero manual annotation — and shows GPT-6 Astra falling from perfect depth-1 repairs to 64% when eight dependent behaviors must be restored together.

Clone the App, Not the Code: Microsoft's ProgramDistill Turns Working Software Into 4,063 Verifiable Coding Tasks

Every coding-agent benchmark has the same dirty secret: somebody had to write the tasks. Human annotators dream up a bug, plant it in a repository, write a patch, and pray that the “gold” solution is actually gold. The result is a corpus of a few hundred hand-crafted puzzles that saturate within months — SWE-bench Verified’s 500 tasks are now routinely cleared by frontier agents. On September 14, Microsoft Research’s Froggy team in Montréal (with KAIST research intern Jeonghye Kim) published ProgramDistill, a pipeline that flips the whole arrangement: instead of humans inventing tasks for software, working software generates its own tasks.

The numbers alone justify attention. From 26 open-source web applications, ProgramDistill’s multi-agent pipeline mined 1,975 replay-verified behaviors and automatically constructed 4,063 coding tasks — 2,862 atomic (restore one missing behavior) and 1,201 cumulative (restore several chained behaviors) — with no manual annotation at any step. Every task carries its own machine-checkable verifier, because the verifier is simply the application’s own recorded behavior.

How the trick works

ProgramDistill is built around what the authors call a mine-craft-patch pipeline, orchestrated by multiple LLM agents. It starts from an application that actually works — with source code available to the pipeline, though never to the agent being tested.

Mining agents explore the application and record replayable traces. A trace stores browser actions, expected success signals, and an optional prerequisite link to an earlier trace. Because the actions can be replayed mechanically and the expected signals checked without any LLM in the loop, each trace doubles as an executable behavioral verifier. Verified traces then seed the discovery of dependent behaviors — a card must exist before it can be dragged, a board must exist before it can hold cards — and only traces that reproduce reliably are kept. Across the 26 applications, these prerequisite relationships form branching dependency trees, with one StreamView lineage reaching depth 12.

Crafting agents remove the implementation of selected behaviors. The pipeline surgically deletes the code behind a chosen behavior, then runs build and replay checks to confirm three things: the prerequisites still work, the target behavior now fails, and a stored gold patch restores it. That last check is the one most synthetic-task generators skip — ProgramDistill proves every task is solvable before shipping it.

Coding agents then repair the damaged application by interacting with a working reference version through the browser, with no access to the reference’s source code or the gold patch. The same replay mechanism that validated the mask now grades the repair.

The design quietly solves the contamination problem that plagues synthetic benchmarks. The reference apps are real, popular repositories, but the masks — which behaviors were removed, and in what combinations — are generated fresh by the pipeline, and an agent cannot memorize an answer it has to discover by clicking through a live app.

What the benchmark found

The headline evaluation, ProgramDistill-300, pits nine coding agents against 300 tasks across 26 applications at restoration depths 1 through 8, where depth counts how many dependent behaviors must be restored together under binary scoring.

At depth 1, GPT-6 Astra solves every task — a reminder that single-behavior repair, the thing most benchmarks test, is approaching solved. But success drops to 64.0% at depth 8, where later behaviors depend on state established by earlier ones. Every other evaluated agent retains less than half of its depth-1 performance at depth 8.

The full-reconstruction setting is harsher. Given a minimal scaffold, a product-level description, and browser access to the reference, agents must rebuild entire applications, evaluated on 590 individual behaviors and 413 cumulative workflows across 12 apps. The results separate the frontier sharply:

AgentIndividual behaviorsCumulative workflows
GPT-6 Astra (max)58.98%49.15%
Claude Opus 5 (max)42.03%28.81%
GPT-5.6 Sol (max)33.39%21.07%

Even the strongest agent recovers only 49.2% of cumulative workflows — implementing useful pieces, it turns out, does not guarantee the pieces work together. These reconstruction runs average roughly 700 agent steps, peaking at 1,921.

The behavioral profile of a winner

The most strategically interesting finding is how the strongest agent works, not just how much it scores. GPT-6 Astra is what the authors call observation-intensive and edit-light: it averages 96.3 current-app observations per trajectory — roughly twice GPT-5.6 Sol’s 45.8 — while making the fewest edit/write steps of any agent, just 9.9 per trajectory. Its winning pattern is extensive behavioral checking of its own implementation, followed by highly selective code changes. The trajectory reads like iterative development against a live reference, not one-shot code generation.

The failure analysis quantifies the alternative’s cost. Across 977 failed atomic behaviors in 36 reconstruction runs, 59.2% involved behaviors the agent never observed in the reference at all. The remainder split into 27.9% producing the wrong state, route, or result; 11.1% with the wrong observable form; and 1.8% observed but simply never implemented. Exploring the reference is the bottleneck.

One documented case is almost poetic: an agent implements card dragging, the card visibly moves, and the feature looks done — but after the same drop gesture, the resulting card order differs from the reference. The interaction works; the resulting state is wrong. The agent validated other board interactions after its final edit instead of replaying the exact workflow that would have exposed the mismatch.

There is also an effort-allocation pathology worth flagging for anyone building agents. From depth 1 to depth 8, the code that must be restored grows by more than 9× and the actions across target traces grow by more than 10× — but agent observation and editing effort per repair target declines as tasks deepen, with observation shrinking most sharply. Agents collectively under-invest in exactly the checking that deeper workflows demand, and their final patches leave more of the masked implementation unrestored. More browser feedback doesn’t fix it either: Claude Opus 5 receives substantially more browser-state feedback than Astra yet recovers fewer behaviors. What matters is whether the agent identifies the right details, implements them faithfully, and rechecks the relevant workflow after changing it.

Why it matters

ProgramDistill arrives at a moment when the industry is drowning in benchmark inflation. By making real applications the task generator and replay the judge, Microsoft has produced a yardstick that is simultaneously scalable, contamination-resistant, and brutally honest about compound workflows — the shape of actual product development, where features arrive as interdependent chains rather than isolated diffs.

The roadmap the authors sketch points in two directions: using restoration depth as a curriculum for training reference-guided agents (with replay-based rewards for reinforcement learning), and extending the paradigm to multimodal agents that interact with references via screenshots and are graded on visual fidelity alongside behavioral correctness.

For developers and enterprises betting on coding agents, the practical takeaway is sobering: single-feature repair is largely solved, but ask an agent to rebuild or deeply restore a stateful application and even the best current model delivers barely half of the workflows. The distance between “can fix the bug” and “can rebuild the product” is now a measured number — and it is 49.2%.