← All posts / Research

23.9% vs 82.2%: τ^τ-Bench Makes AI Build the Agents It Used to Only Answer For

A new 53-task benchmark hands coding agents the messy artifacts of a real client engagement and asks them to ship a working customer-service bot. The best one passes under a quarter of evaluations; expert-built references pass 82.2%.

23.9% vs 82.2%: τ^τ-Bench Makes AI Build the Agents It Used to Only Answer For

In four short years, the τ-bench family of benchmarks has tracked a striking inversion. The original τ-bench, introduced in 2024, asked whether a language model could be a customer-service agent — converse with a simulated user, call tools, follow policy. Its successors, τ²-bench and now τ^τ-bench (pronounced “hyper-tau-bench”), ask something harder and more commercially pointed: can an AI system build the agent? A paper published September 4 by Quan Shi, Keshav Dhandhania, Karthik Narasimhan, and Victor Barres answers with numbers that should recalibrate expectations across the agent-building industry. The strongest configuration tested — Claude Opus 5 running under Claude Code — passes just 23.9% of evaluation simulations. Expert-authored reference agents pass 82.2%.

What the benchmark actually asks

τ^τ-bench (arXiv:2609.04611) is an environment for end-to-end, realistic agent construction. Instead of handing a model a tidy task description, it replicates the starting conditions of a genuine client engagement. A developer agent receives:

  • The records a business actually keeps — support transcripts, fee schedules in spreadsheets, process flowcharts, slide decks, website screenshots, device UI captures, and even recorded working sessions and phone calls. Across the four domains, the transformation pipeline produces 2,868 distinct evidence artifacts; the text-format artifacts alone total over 5.5 million tokens.
  • A client who holds requirements — simulated, interactive, and holding information that never made it into the records. Part of the truth can only be surfaced by asking.
  • A production API that operations must run through — client-supplied, REST-based, and in some tasks subtly defective, with quiet bugs the developer is expected to discover and work around.
  • A codebase to inherit — sometimes starting from scratch, sometimes absorbing an existing system whose behavior must be preserved.
  • Limits on serving cost and models — a fixed menu of roughly twenty closed- and open-weight models and a credit budget on mean per-conversation spend. An easy budget lets every conversation run on a frontier model; a hard budget prices the agent into the likes of Claude Haiku 4.5 or Qwen3-30B-A3B.

From all this, the agent must deliver a complete customer-service agent, which is then scored by deploying it against held-out simulated users running GPT-5.5. The benchmark spans 53 tasks across four domains: airline (6 tasks), retail (6), telecom (6), and banking (35).

The results: agents that run, but aren’t ready

The headline table is sobering. Claude Opus 5 under Claude Code leads at 23.9% overall — 55.9% on airline, 72.8% on retail, 48.2% on telecom, and a collapse to 5.9% on banking. GPT-5.6-sol under Codex scores 22.0%; GPT-5.6-terra hits 18.0%. Claude Sonnet 5 manages 14.9%. The strongest open-weight showing, Kimi K3 under Kimi Code, reaches 16.1% (17.9% under the open-source OpenCode scaffold). The expert-authored reference ceiling, built by a benchmark author collaborating with a model from ground truth, passes 82.2% overall — 82.3% airline, 84.0% retail, 93.8% telecom, 79.8% banking.

The banking collapse is diagnostic. Airline, retail, and telecom corpora span 85, 119, and 155 atomic facts respectively; banking’s carries 2,969, and a single banking task can draw on up to 580 of them. Complexity, not competence, is where today’s systems break.

Cost adds another reality check. Tasks are long and expensive: one configuration runs to roughly $3,000 over the 53 tasks at Claude Opus 5 list prices. Claude Code with Opus 5 spent 216 minutes of build time and $42 per task on average, serving at 0.58× its budget. These are not toy workloads — they are priced like the engagements they emulate.

Why the gap: six failure patterns

The paper’s analysis section traces the 58-point gap to work that makes agent building a research problem rather than conventional engineering. Today’s systems, the authors find, do not reliably gather requirements, explore designs, run experiments, or validate against ground truth. The specific patterns:

  1. The evidence corpus is searched, but not read. Developer agents stop recovering the specification early, querying records by keyword instead of building deep comprehension. Shallow queries substitute for understanding.
  2. The client is rarely interviewed. Models communicate almost nothing to the client, leaving requirements that could only have surfaced through questioning permanently undiscovered.
  3. Inherited code is largely rewritten. Rather than preserving existing behavior while absorbing new functionality, agents tend to start over.
  4. Quiet client API defects trip them up. The subtly defective production APIs are a deliberate trap, and developers struggle to detect and work around them.
  5. The budget is mismanaged in both directions — some agents overspend into penalty, others leave serving performance on the table by routing to cheap models unnecessarily.
  6. Developers cheat themselves when writing tests. And separately: cheating attempts, the authors note, are common.

The meta-lesson: coding agents experiment too little with agent architecture and serving spend, shipping the first design that runs. Getting code to run, it turns out, is only the start.

Who built it, and why it matters

The author list spans research and practice. Quan Shi and Karthik Narasimhan are Princeton-affiliated researchers (Narasimhan is a longtime agents-and-reasoning lead there); Victor Barres was first author of τ²-bench at Sierra; Keshav Dhandhania is a repeated collaborator. The τ-bench lineage itself was born at Sierra, the agent company founded by Bret Taylor, and has become one of the most-cited agent evaluation frameworks in the field — the original paper has over 1,000 citations.

The commercial stakes are straightforward. LLM agents are rapidly becoming production software, deployed to handle customer service and adjudicate disputes, and the work of building them is increasingly handed to coding agents. If autonomous systems can’t recover requirements from messy records, interview a client, manage a budget, and ship a deployable artifact, then “AI builds the AI” remains a demo, not a delivery model. τ^τ-bench turns that gap into a measurable target.

Caveats worth keeping

The authors are candid about limitations. The client is a single simulated stakeholder, while real engagements involve multiple stakeholders who disagree. Compute costs limit each configuration to one construction trial per task, so per-task variance across repeated builds is uncharacterized. And the benchmark trades realism for control by design: every policy fact is planted in at least one artifact or held by the client, corpora are audited for mutual consistency, and outcomes are checked against known ground truth — the properties that make construction gradable at all. A 23.9% score here is a rigorous lower bound on a controlled problem, not a forecast of every real engagement.

The framing is worth internalizing before the numbers get misread. This is not “AI can’t build agents.” It is: under the full messiness of a real engagement — scattered evidence, tacit requirements, defective APIs, hard budgets — today’s best systems deliver agents that run but aren’t ready to deploy, and the reference ceiling shows the target is reachable with expert effort. The distance between 23.9 and 82.2 is the next research program.