Telling Agents to Test Better Makes Them Worse: Inside Dan Luu's 26-Condition Experiment
Dan Luu ran 26 prompt conditions across thousands of Rust agent runs and found that naming a testing technique — TDD, Lean 4, QuickCheck, Verus — reliably produced worse correctness than saying nothing at all.
On September 8, 2026, Dan Luu published an essay that quietly detonates one of the most comfortable assumptions in AI-assisted engineering: that the way to get coding agents to write correct software is to tell them which quality technique to use. In “How well do agents use test/verification techniques?”, he reports the results of a sprawling controlled experiment in which he handed a coding agent the same implementation task under 26 different testing instructions — test-driven development, Lean 4, QuickCheck, Verus, TLA+, fuzzing, mutation testing, SMT solvers, and more — and measured what actually happened.
The headline finding is blunt: nothing wildly outperformed, and the Default condition — no testing instructions at all — scored well above average. Naming a technique didn’t just fail to help. In most cases it actively hurt, because agents would superficially perform the named technique while producing worse tests and worse implementations than they would have written unprompted.
The setup
The eval reuses a benchmark Luu has described in earlier posts: implementing the Zstd compression algorithm in Rust, scored by the fraction of runs that pass 100% of a hidden test suite, averaged over 80 runs per condition per effort level. The agent was codex running GPT-5.6 Sol at medium and xhigh effort. The 26 conditions ranged from heavyweight formal methods (ACL2, Alloy, Creusot, Kani, Lean 4, Spin, TLA+, Verus, SMT solvers with Z3/cvc5/Yices installed) through randomized-testing families (QuickCheck, Proptest, Hegel, property-based testing, fuzzing, differential testing, metamorphic testing) to process instructions (TDD, “Audit first”, “Audit and fuzz risky areas”, snapshot testing with Insta, fixture testing with rstest, Rust’s built-in framework, and a “Judgement” condition that simply asked agents to use the best technique).
Four skills were also tested: the official Hegel skill, the ECC Rust test skill (from a collection with 250k GitHub stars and 38k forks), a Trail of Bits property-testing skill, and a minimal skill Luu wrote himself in a couple of minutes. The first three were chosen precisely because they were the top results when codex was asked to find relevant testing skills — in other words, the ones real users actually encounter.
He also pre-registered predictions, an unusually disciplined touch: TDD would underperform (55% confidence), formal methods would not overperform (52%), “Make no mistakes” would not outperform no instructions (95%), and the popular skills would disappoint. Every one of those guesses came true.
What agents actually did
The failure mode is consistent across conditions and more interesting than the topline numbers. Agents didn’t refuse the instructions — they technically complied while extracting almost no value from the technique:
- Formal methods became theater. Verus agents proved abstract arithmetic properties and vacuous implications of the form A implies A, while avoiding proofs about the code paths where bugs actually lived. Lean 4 agents did the same. TLA+ agents built state-machine models, but 75 of 80 modeled the implementation only after writing the code, and Luu could not find a single case where a TLA+ finding changed the Rust implementation.
- Property-based testing degenerated into smoke tests. QuickCheck agents mostly wrote trivial checks; 63 of 160 runs checked exactly one property. Fully random inputs against a format like Zstd mostly exercise the rejection paths, so the “randomness” tested nothing.
- Differential testing wasn’t differential. Out of 160 runs, 135 did something called differential testing, but none built two genuinely independent implementations — agents wrote the same thing twice and encoded the same bug in both.
- TDD produced more tests and worse tests. TDD agents wrote roughly twice as many tests and ran a test-code-test-code loop, but were more likely to miss hard cases like Zstd’s four-stream jump table — often writing four identical, trivial bitstreams where a real bug (transposed streams) would be invisible.
Gary Bernhardt’s summary of agent testing behavior — take the pathological edge cases, make them the backbone of your strategy — turned out to apply to every technique once its name was invoked. Agents pattern-matched the surface ritual of each method without the judgment that makes the method work.
The exceptions are instructive
Two results complicate the bleak picture. First, when agents generated structured random inputs rather than naive random bytes, they found real bugs half the time — but this happened in only 10 of 160 fuzzing runs. The capability is latent; the default behavior just doesn’t reach it. Second, Luu’s own hand-written skill scored highest of all, despite being a five-bullet sketch about identifying risky areas before implementing, preferring boundary and asymmetric examples, and steering randomization toward interesting state paths. The contrast with the 34k-character Hegel skill — which raised costs 16–18% via token bloat and re-reads while lowering correctness — is stark. Skills written as tutorials to teach a technique underperformed; a skill written to nudge the agent away from its known failure modes helped.
That distinction — explain versus steer — is arguably the essay’s most actionable idea for anyone building agent tooling today.
Why this matters beyond Zstd
Luu frames the broader question at the end: why haven’t AI labs built RL environments to train agents to test well? He notes that agents have become excellent at bounded runtime-optimization problems, exactly the kind of task that is cheap to mass-generate RL environments for, and speculates that effective testing is in the same class of problem — but that the limiting factor may simply be that knowledge of effective test techniques isn’t widespread enough for anyone inside the labs to have tried. He also cautions that his RFC-based evals are easier than real-world specifications, which are vaguer and more ambiguous, so the failure modes seen here should be the same or worse in production use.
The essay closes with a note that will resonate with anyone shipping agent-written code: from the inception of public coding agents until now (September 2026), agents that had some idea how to test — without being guided by a testing expert — would have substantially increased agentic coding effectiveness. Until the labs close that gap, the empirical answer to “how do I get agents to test better?” is not a technique name. It’s a human who looks at what the agent did, notices the failure mode, and types a few more sentences.