← All posts / Research

Write the Planner, Freeze It, Test It: Coding Agents Beat Hand-Engineered Robot Planners at Their Own Game

A new paper shows Claude Code and Codex agents can synthesize generalized task-and-motion-planning programs that outperform classical planners 56–95% vs 47% across 98,000 evaluation episodes.

Write the Planner, Freeze It, Test It: Coding Agents Beat Hand-Engineered Robot Planners at Their Own Game

For roughly a decade, task and motion planning (TAMP) has been one of robotics’ most stubborn problems. The challenge is deceptively simple to state: a robot must decide both what to do (a discrete sequence of actions) and how to do it (continuous geometric, kinematic, and dynamic motions that respect physics). A pick-and-place task might require choosing an order of grasps while simultaneously guaranteeing that an arm never collides with the table. The discrete and continuous pieces are tightly coupled, and solving them jointly has traditionally demanded deeply specialized engineering — domain-specific planners like PDDLStream, hand-tuned samplers, and years of expert iteration.

A new paper posted to arXiv on September 24 asks a blunt question: do we still need any of that hand-engineering? In “Coding Agents for Generalized Task and Motion Planning Problems,” researchers led by Matteo Merler and Tom Silver evaluate Claude Code (Opus 5) and Codex (running both GPT-5.6 Sol and GPT-6 Astra) as generalized TAMP planners — and find that the agents convincingly outperform the classical stack.

The setup: programs, not policies

The experimental design is what makes the paper credible. Each coding agent receives a task description and access to a simulator, then works within a fixed synthesis budget — choosing freely how to interact with the environment, run experiments, and iterate — until it produces a program. That program is then frozen and evaluated on unseen problem instances it never encountered during development.

This “write it, freeze it, test it” protocol closes the loopholes that usually inflate agentic benchmarks. The agent cannot adapt at evaluation time, cannot memorize test cases, and cannot hack a reward signal that only exists at training. Generalization has to live inside the synthesized code itself.

The evaluation surface is substantial: 28 simulated environments drawn from the KinDER and PDDLStream benchmark suites, with object counts pushed beyond what the original benchmarks were designed for. Across all program synthesis methods, the team evaluated 980 generated programs on 100 held-out instances each — 98,000 evaluation episodes in total.

The results

The headline numbers: all three agent configurations beat the hand-engineered planners in mean success rate, scoring 56% to 95% versus 47% for the classical planners on the 16 environments where a planner baseline is available. The agents also outperformed both one-shot generation (writing the program in a single pass with no environment interaction) and an LLM-based generalized planning baseline.

Two secondary findings matter as much as the top-line score:

  • Scaling behavior. As object counts grow beyond the original benchmark regimes, the agents’ programs maintain higher success rates than the classical planner. TAMP solvers are notorious for combinatorial blowup as scenes get cluttered; the synthesized programs degrade more gracefully.
  • Compute efficiency. At evaluation time, the agent-written programs use roughly an order of magnitude less computation per instance than the planner. For real robots, where planning latency translates directly into how fast an arm can move, that margin is the difference between a demo and a product.

Perhaps the most interesting evidence is in the logs. The researchers report that agents used their interaction budget to calibrate physical models (refining friction or grasp parameters against simulator feedback), test edge cases (degenerate object arrangements, near-collision configurations), and refine strategies when initial attempts failed. In other words, the agents behaved less like text generators and more like experimentalists — running the same hypothesis-test-debug loop a robotics graduate student would run, compressed into an automated synthesis budget.

Why this lands now

The result arrives at a moment when the robotics industry is publicly struggling with exactly this problem. Tesla’s Optimus program, as reported this week, is building hundreds of robots per week yet still finds that its AI “can’t yet reliably handle a wide range of tasks,” with basic behaviors taking days to learn and fine manipulation bottlenecked by hardware. The gap between impressive hardware volume and fragile task generalization is the generalized-planning gap this paper attacks.

The economics also explain the timing. Classical TAMP expertise is scarce and expensive — a single domain can consume a doctoral thesis. Coding agents rent that expertise by the hour. If an agent can synthesize a competent, reusable planning program for a new environment in a bounded budget, the marginal cost of adding a new task class collapses. The paper positions coding agents as “a strong baseline for generalized TAMP” — baseline being the operative word: the thing every future approach must now be compared against.

The caveats worth keeping

Honest framing requires noting what this paper does not show. All 28 environments are simulations with full observability and object-centric state representations — the cleanest possible setting for TAMP, and far from the partial-observability, sensor-noisy reality of a warehouse floor. The 56%–95% spread across configurations means model choice and interaction strategy still matter enormously; the lower bound leaves nearly half of instances unsolved. Sim-to-real transfer of synthesized programs remains untested. And the synthesis budget itself, while fixed, is not free — the evaluation-time efficiency does not account for what the agent spent learning the environment.

But these are the normal caveats of a strong empirical result, not disqualifiers. The authors release all code, including the full prompts given to the agents — an unusually reproducible posture for a field where agentic pipelines are often described but never shipped.

The bigger picture

The deeper significance is architectural. Robotics has oscillated between two poles: hand-engineered symbolic planning (interpretable, brittle) and end-to-end learned policies (flexible, opaque). This paper points at a third path — agent-synthesized symbolic artifacts. The agent does the slow, expensive exploration offline; the deployed robot runs a fast, inspectable, frozen program. You get the interpretability of classical planning with the development velocity of modern LLMs.

For the coding-agent ecosystem, it is one more domain — after competition-level mathematics, formal verification, and nine-loop physics calculations — where frontier models dropped into an agentic harness produce artifacts that outperform human experts’ handiwork. The pattern is becoming hard to ignore: the scarce resource is no longer the ability to write a solver. It’s knowing which problems are now solvable by asking.