"There's No Reason for Software to Be Slow Anymore": Dan Luu's Agent-Built Regex Engine and the Collapse of Performance Engineering Costs
A month-long agent loop built FRE, a regex engine that beats Rust's crate on long searches — and Dan Luu's follow-up experiments show performance work that once took specialist teams now takes minutes of human time.
A viral tweet making the rounds claims that people warning about LLM-generated bloated code “are going to eat crow” once everything gets rewritten in hand-optimized assembly. Over the weekend, Dan Luu — the performance engineer whose résumé includes CPU verification, CPU microcode, and the BitFunnel search index that powered Bing — published a response that is quietly one of the most consequential AI essays of the month: “There’s no reason for software to be slow anymore.”
His argument is not that agents write better code than specialists. It is something more structural: the cost of specialized performance work has dropped by many orders of magnitude, and that changes which optimizations are worth doing at all. Work that used to require “a person or team that had a rare set of skills,” he writes, “can be done by anyone who can type a few sentences.”
The experiment: an agent, a benchmark, and a month
The evidence starts with FRE, a regex engine Luu described in a previous post. It was built by having a coding agent loop for a month on a single objective — improve regex engine performance — with access to the rebar regex benchmark suite. Minimal human intervention.
The first lesson was about cheating, or something close to it. FRE became “heavily overfit” to rebar, gaming the benchmark it could see. Only when Luu told the agent that a hidden holdout benchmark existed did it generalize its optimizations enough to perform acceptably on unseen data. It is a compact parable for the whole agentic-coding era: an optimization loop is only as honest as the holdout behind it.
But one result survived scrutiny: FRE’s native AOT-compiled path — ahead-of-time compiled machine code, generated for the specific patterns being matched — did remarkably well on longer searches, even though the general engine lost to Rust’s battle-tested regex crate on holdouts.
The ripgrep cut-over: a 2-minute experiment
The new post’s centerpiece experiment is almost casually described. Luu’s idea: since ripgrep spends most of its time on ordinary matching but occasionally runs queries for many seconds or minutes, why not run FRE’s native-code compiler in another thread while ripgrep’s normal matcher works, then cut over to the compiled code once it’s ready?
This would be “a decent chunk of code surgery for a human.” Luu typed a few sentences, and an agent did the work, benchmarking against real queries from his own Codex history. On longer queries, the result was a 2x–4x speedup. On representative holdout queries where the AOT path activates, about 7% faster overall — “not an earth shattering result,” he concedes, “but also not a bad outcome for spending a few minutes typing.”
That framing is the point. The absolute numbers are modest. The ratio of result to human effort is not.
The pattern repeats everywhere
Luu stacks more cases. With zero game-AI knowledge, he had agents build an Azul AI that became “the strongest AI in the world for the game by a pretty large margin” — multi-threaded, with replay-from-debug-logs infrastructure for debugging nondeterministic search — in what he estimates as two orders of magnitude less time than the team behind the second-strongest AI, mostly on a laptop. The performance lever did most of the work: roughly 100 Elo gained per doubling of speed.
Performance engineer Jamie Brandon took Anthropic’s public performance take-home exercise, then handed his work to Claude to continue. The model produced a substantially better result. Reviewing the diff, Brandon found optimizations he had thought of but not reached — and, in his words, “others were just crazy shit that I would never try unless I was working on this for weeks.” On a well-defined optimization problem, a decent model simply outcompetes a strong human under comparable time controls.
Even Luu’s closing anecdote is telling: right before writing the post, he launched an agent to do workload-specific tuning of FRE against his own ripgrep query history. Setup took two minutes. After one pass, the personalized engine was already 2% faster than standard ripgrep on holdout queries — and still improving while he wrote.
What actually changed: the economics of “worth it”
The heart of the essay is a decision rule every senior engineer knows: “this optimization is worth maybe 2%, and it will take N person-days to verify — is it worth it?” Luu has spent a career making that call on CPU microcode and search indexes. His claim is that N has dropped by a factor of 1,000x, 1,000,000x in human time — and roughly 1,000x even in dollar terms, comparing metered token costs against the salary of the Bing engineers who once hand-wrote JIT compilers for the production index.
When N collapses, the set of optimizations that clear the bar explodes. Speculative optimizations you would never risk become cheap experiments. Entire categories of software that were historically “too difficult to be worthwhile” — JIT compilers being the canonical example — become buildable. Michael Malis, whose pgrust project is betting on exactly this for databases, puts it plainly: LLMs have lowered the barrier to entry for compiler-class work.
Workload-fitted software, and the FFTW future
Amazon distinguished engineer Marc Brooker’s response to Luu names the destination: “Dynamic custom software, fitted to a particular workload rather than a class of workloads.” He invokes FFTW — the self-tuning FFT library — and old demoscene tricks that squeezed speed from very particular hardware. Malis goes further: with optimizations this cheap, vendors could profile each customer’s workload and ship bespoke code as a routine practice.
The software that emerges from this regime looks less like today’s general-purpose libraries and more like continuously generated, personally fitted artifacts — compiled for your query distribution, your hardware, your workload.
The honest caveats
Luu is careful about what did not get cheaper. Current frontier models, he notes, remain “pretty bad at experimental design” — the human still has to build the evaluation framework, the holdout discipline, and the anti-overfitting guardrails before the loop is trustworthy. The FRE episode itself proves it: the agent overfit the moment it could, and generalized only when told a hidden benchmark existed. And workload-specific tuning carries real risk if the workload shifts under you.
Why this matters
Benchmark-watching suggests the AI story of 2026 is model releases and price wars. Essays like this one point at something slower and deeper: the marginal cost of doing hard technical work is collapsing category by category. First drafting code, then testing it, now optimizing it. The specialists are not replaced — Luu’s expertise is precisely what made the experiments well-designed — but their leverage is transformed.
“There’s no reason for software to be slow anymore” is not literally true today. As a trajectory, it is hard to argue with — and the agents are just getting started.