← All posts / Tools

Science, Autotuned: Claude Code Learns to Build Evals and Hillclimb Its Own Agents

Anthropic's new /claude-api build-eval and hillclimb workflow turns agent optimization into a controlled experiment, with train/test splits, noise floors, and automatic rollbacks — one demo cut support costs 80% while raising accuracy.

Science, Autotuned: Claude Code Learns to Build Evals and Hillclimb Its Own Agents

There is an awkward problem at the heart of “improving” an AI agent: sometimes you make the benchmark better without making the agent better. Change the prompt. Swap the model. Add a tool. Crank up the reasoning effort. Run the test again. The score goes up. Ship it? Maybe — or maybe you just taught your system to pass that particular test.

On September 28–29, 2026, Anthropic shipped a disciplined answer to that question inside Claude Code. A new auto-tuning workflow for the official claude-api skill gives Claude two jobs: first, help you build an evaluation that actually resembles the work your agent does in production; second, hillclimb your system against that evaluation — repeatedly modifying prompts, models, and settings while checking whether each “improvement” survives on examples it was never allowed to study. Lance Martin, the Anthropic engineer who wrote the underlying guide, walked through the method as the news broke, and The Neuron published a detailed explainer on September 29.

The practical core is almost painfully simple: real tasks → reliable grader → baseline → one change → test it → keep or revert. Which sounds suspiciously like “do science,” except Claude is now also the lab assistant.

Why this is harder than it sounds

An eval is a test suite for AI behavior. Normal software tests are binary: did the function return the right number? Agents are messier. They write emails, investigate incidents, route support tickets, browse websites, and make judgment calls where two different outputs can both be correct — and the same model can answer the same task differently across runs.

Anthropic’s workflow insists an eval needs four properties before you optimize anything:

  1. Tasks should look like production. If customers ask your support agent about refunds, billing, and account access, the eval should test those — not a pile of toy questions.
  2. Better models should score better. If your expensive frontier model underperforms a smaller one, something is wrong with the task, grader, or environment.
  3. There must be room to improve. Above roughly 95%, the tooling itself warns you and suggests optimizing cost or latency instead.
  4. Scores must be stable. If the same system swings between 65% and 90% across runs, you’re measuring noise, not product quality.

That last point matters most once you let an AI optimize the AI. If your ruler changes length every time you measure, an automated loop will happily spend its time getting extremely good at moving the ruler.

The sneakiest trap: benchmarking your model’s potholes

Anthropic calls out a failure mode every team eventually discovers the hard way. Run your agent 1,000 times, collect the 50 ugliest failures, and turn those into your benchmark. Sounds smart — until you remember that modern models have a “jagged capability surface,” strangely brilliant at one task and strangely terrible at a near-identical one. A benchmark selected purely from one model’s failures measures the quirks of that model. When the next model arrives, your carefully curated benchmark becomes a museum exhibit.

The /claude-api build-eval command tries to automate the boring part safely. It interviews you about what you’re evaluating, assembles test cases with a clear evidence hierarchy — production transcripts first, then bug reports and support tickets, then five to ten hand-written examples, and only then synthetic cases generated from the codebase — and renders a local page where a human approves the inputs before anything counts. The grader gets equal scrutiny: constrained tasks get programmatic grading (exact match, valid JSON, passing tests), open-ended tasks get an LLM judge with a rubric of specific checkable claims — and Anthropic explicitly recommends never letting the model under evaluation grade itself. Claude even runs the grader twice on identical output to detect verdict-flipping before you trust it.

Hillclimbing with a held-out set

Once the eval checks out, /claude-api hillclimb lets Claude start changing the system. You decide which surfaces it may touch: system prompts, skills, tool descriptions, model choice, reasoning effort, other API parameters, or the surrounding agent harness. You also declare what you care about — maximum accuracy, lower cost, or similar performance from a cheaper model.

Then comes the detail that is the entire game: Claude randomly splits the evaluation into a train set and a held-out test set. It may inspect failures from the training examples. It never sees the held-out answers while deciding what to change. Each round follows the same loop — inspect train failures, find a root cause, propose one patch, rerun the eval, compare train and test, keep or revert. One change per round, so you can attribute causality. Train score up but held-out flat? That’s flagged as overfitting and reverted. Before any of it starts, the workflow measures the eval’s noise floor and checks whether it’s smaller than the smallest improvement you’d actually care about; if not, it asks for more examples instead of pretending to see signal.

And when progress stalls after two or three rounds, Claude stops randomly poking things, groups the remaining failures by root cause, and asks whether the problem is the agent, the instructions, the grader, or the benchmark itself.

The numbers that sell it

The customer-support demo is the one worth writing down. An internal benchmark of 44 tickets: 30 for the hillclimb search, 14 held out and never shown to the optimizer. The starting configuration — a heavyweight model at high effort with a prompt full of ritual — scored 74.4% at 4.6 cents per ticket. Claude first audited away mandatory tool-call rituals, a scratchpad step, and contradictory rules. It then tried Opus 5.5 at low effort (87.8% at ~1.9 cents), stepped down to Sonnet 5 at low effort (88.9% at ~1 cent), and finally improved the routing prompt with a refund-cap cross-reference.

The headline isn’t the 98.9% it eventually reached on the training tickets — Claude had studied those. It’s the held-out 14: the original setup scored 78.6%; the final setup scored 90.5% at roughly one-fifth the cost. More accurate on unseen work, 80% cheaper.

The second example is arguably more instructive. Hillclimbing Anthropic’s own claude-api skill from ~66%, Claude found eight missing API features (74%), fixed errors in the C# and Java type tables (77%), and — after progress stalled — noticed Claude kept writing older API patterns remembered from pretraining. A mapping table from old patterns to the current API pushed it to 80%. Then came the twist: some stubborn “failures” turned out to be bad evals. One grader expected an error chain the task never asked for; another contradicted Anthropic’s own documentation, and testing the real API proved the docs right. Fixing the benchmark along with the skill reached ~88%. Sometimes the agent is wrong, sometimes the instructions are wrong, sometimes the test is wrong — a serious workflow needs to be able to discover all three.

Limits worth respecting

Automated hillclimbing does not solve the hardest problem: Claude can only optimize the objective you give it. Reward the wrong behavior and you’ll get better at the wrong behavior, faster. And when a production agent depends on live Slack messages, changing memories, or external websites, an offline held-out set may still fail to represent reality — teams running long-lived agents in production report exactly this friction with offline evals. Anthropic’s workflow reduces one class of self-deception; it doesn’t make measurement magically objective.

But the direction is unmistakable. Agent development is shifting from prompt craft — clearer instructions, more examples, better models — toward measurement as infrastructure: change the system, evaluate real tasks, quantify uncertainty, test unseen cases, automatically revert regressions. Anthropic’s own changelog shows the tooling iterating quickly, with hillclimb already skipping prompt rewordings too small for the eval to measure. When your AI coding tool can finish a run with “I changed nothing; that was the correct decision” — and mean it — evals have stopped being a research accessory. They’ve become part of the product.