← All posts / Research

HarnessTax: Your Coding Agent's Model Is Fine — the Wrapper Is Costing You 2x

Berkeley and Arena measured 21 model–harness pairs and found harness choice barely moves success rates but can multiply token costs — Claude Code ran ~2x Pi's bill for a 1.1-point gain.

HarnessTax: Your Coding Agent's Model Is Fine — the Wrapper Is Costing You 2x

Every coding-agent comparison you have read probably focused on the model. Which frontier model solves the most GitHub issues? Which one writes the cleanest diff? A new study from UC Berkeley and Arena flips that question on its head: when the model is held constant, the harness — the software system that manages the model’s tools, context, and execution loop — has almost no effect on whether tasks get solved, but a large effect on what you pay. The team calls the difference a “harness tax,” and for teams running agents at scale, it may be the most expensive line item nobody benchmarks.

What the study did

The researchers evaluated 21 model–harness pairs: seven models across three harnesses — Claude Code, Codex CLI, and Pi. The model lineup spans the current commercial spectrum: Claude Fable 5, Claude Opus 4.8, Claude Sonnet 4.6, Claude Haiku 4.5, GPT-5.6 Sol, GPT-5.6 Luna, and the open-weight Kimi K3.

Two open-source benchmarks anchored the evaluation: SWE-bench Lite, which measures software modification capability on real GitHub issues, and Terminal-Bench 2.0, which evaluates complex command-line tasks. The methodology was deliberately tight: the same 30 randomly sampled tasks from each benchmark, three runs per task to capture variance, each harness at its high-effort setting, and attempts capped at 100 agent turns. Success was judged by each benchmark’s official evaluator — no homegrown grading — and token costs were computed from a fixed direct-API price list dated September 1, 2026, applied identically across harnesses. The team also took care to neutralize environmental differences: for SWE-bench Lite, external network access was blocked in all task containers, and web tools in Claude Code and Codex were disabled so that no harness got an unfair lookup advantage.

The statistical treatment is more rigorous than most agent evaluations: 95% confidence intervals were estimated with 10,000 bootstrap resamples, and the results are positioned against Pareto frontiers — the staircase of best success rates achievable at or below a given cost — so the reader can see cost-quality trade-offs rather than a single leaderboard number.

Finding 1: harnesses tax your wallet, not your success rate

The headline result is stark. Holding the model constant, average success-rate differences across harnesses stayed within ±2% on SWE-bench Lite and within roughly ±5% on Terminal-Bench 2.0 — while costs diverged by up to 5x.

The emblematic case is Claude Fable 5. It solved 97.8% of attempts inside Claude Code, versus 96.7% in Codex and 96.7% in Pi. But Claude Code cost about twice as much as Pi per attempt ($1.33 vs $0.67). The team is careful to note turn counts were nearly identical — 15.3 turns per attempt in Claude Code versus 15.4 in Pi — so the premium is not explained by the agent taking more steps. The same work, the same outcome, double the invoice.

Aggregating across shared models with geometric means, Claude Code cost about 2.0x Pi and 1.6x Codex on SWE-bench Lite, and 1.5x Pi on Terminal-Bench 2.0. If you accepted your coding agent’s default harness without shopping around, the study suggests you may have been quietly paying a tax that never showed up in any success metric — because it doesn’t change the success metric.

Where does the money go? A clue sits in the very first model call. Across all seven models, Claude Code’s mean initial context was over 10x Pi’s, with longer instruction statements and larger tool schemas — roughly 77,000 characters of tool specification and instructions on average before the agent has done anything. Input-token costs scale with that preamble, though the authors caveat that caching, generated tokens, and subsequent calls all shape the final bill, so initial context alone doesn’t explain everything.

Finding 2: a four-tool harness reaches the Pareto frontier

The second finding is more subversive: Pi, a minimal open-source harness that provides just four tools — read, write, edit, and bash — reached the Pareto frontier on both benchmarks. No proprietary scaffolding, no co-training with the model, no elaborate permission choreography. The bare essentials were enough to be cost-competitive with the best commercial stacks.

That result reframes what “harness engineering” buys you. Feature richness did not translate into measurable success gains on these benchmarks — the agent’s competence lives mostly in the model, and the harness’s job is to pipe intelligence to the right tools without burning tokens. It also has a research-policy implication the authors make explicit: researchers can now work with state-of-the-art coding harnesses without access to proprietary systems, because the open ones are already at the frontier.

The authors resist over-claiming. The findings are limited to two open-source benchmarks that models may have encountered during training, and results may differ on other workloads. Harness complexity should be treated as an empirical trade-off, not a virtue.

Finding 3: models don’t need their home harness

The third finding punctures provider pairing. OpenAI, for instance, describes GPT-5-Codex as optimized for software engineering in Codex. Yet across the six Anthropic and OpenAI models and both benchmarks, an alternative harness achieved the highest observed success rate in nine of twelve comparisons.

The pattern repeats at every tier. Sonnet 4.6 solved 68.9% of attempts in Codex versus 66.7% in Claude Code on SWE-bench Lite, at similar cost. GPT-5.6 Sol hit 83.3% in Pi on Terminal-Bench 2.0 versus 78.9% in Codex — at roughly half the cost ($0.42 vs $0.76). Or, as the blog post puts it more bluntly: your Claude models may not need Claude Code. Model capabilities are generalizable and carry over to other harnesses; a shared provider does not guarantee the best pairing.

There is also a quiet victory for open weights in the data: Kimi K3 sat close to the Pareto frontier near GPT-5.6 Sol on SWE-bench Lite, and just below it on Terminal-Bench 2.0 — competitive coding-agent performance from an open-weight model.

Why it matters

The economics of coding agents are usually discussed as a model-pricing story — dollars per million tokens, subscription tiers, batch discounts. HarnessTax shows there is a second, largely invisible pricing layer: the wrapper’s token overhead, which can swing the cost of identical work by 2–5x without moving the success needle. For an individual on a subscription, that’s noise. For a company running thousands of agent-hours per day, it is the difference between one cloud bill and another.

The study’s methodological contribution may outlast its specific numbers: evaluations that vary only the model, holding the harness fixed, systematically miss this cost axis. The authors argue model evaluations should compare the same model’s cost and success across commonly used harnesses — and envision a future harness that adapts as tasks unfold, rather than forcing users to make configuration decisions by folklore.

Caveats worth keeping: 30 tasks per benchmark is a small sample; benchmarks may be contaminated by training exposure; “high effort” settings and 100-turn caps follow each harness’s own definitions; and real development workflows — evolving requirements, human feedback, multi-session tasks — differ from benchmark containers. The authors release their profiling traces publicly, which is the right call for a result this consequential.

The bottom line

Coding agents are, for day-to-day tasks, interfaces to model intelligence. As models get more capable, they may need less of today’s scaffolding — and the scaffolding you don’t need is pure margin for whoever bills you by the token. HarnessTax is a reminder that in the agent stack, the model is the talent, but the harness is the contract. Read the contract.