Self-Improvement for $150: MIT and Sakana AI's SIFT Cuts the Cost of Recursive Agents
An LLM-judge-guided tree search lets a coding agent rewrite itself to 35.1% on Polyglot in five hours on $150 of API credits — a tenth of the compute of prior methods.
For most of the past two years, “recursive self-improvement” has been a phrase reserved for frontier labs with frontier budgets. The Darwin-Gödel Machine (DGM) showed that a coding agent can rewrite its own scaffolding and get better at coding as a result — but making it work consumed thousands of CPU hours and, in some reported configurations, upward of $22,000 in evaluation spend on SWE-Bench. That price tag effectively gated an entire research direction to a handful of well-funded teams.
A paper released in September from MIT and Sakana AI, now getting broad attention this week, argues the gate was never the algorithm — it was the evaluation bill. SIFT (Self-Improvement via Fast Tree-search) keeps the recursive loop but replaces most of the expensive benchmark re-runs with something far cheaper: an LLM judge that ranks candidate self-modifications by simply reading them.
The bottleneck was never the search
The core observation is disarmingly simple. When a self-improving agent proposes a change to itself, prior frameworks estimated whether the change helped by re-running the modified agent on a subset of benchmark tasks. On Polyglot-50, that subset evaluation alone runs roughly $6.00 and 2.6 CPU-hours per candidate. Judge calls — a model reading two patches and picking a winner — cost a few cents.
SIFT’s answer is to flip the economics. Every candidate patch still enters the tree, but a pairwise LLM-as-a-judge signal does the front-line filtering. Win-loss records are aggregated with a regularized Bradley-Terry model into per-node strength scores, and those scores drive rank-based parent sampling inside a lightweight, fully disaggregated tree search. Expensive downstream task evaluations are reserved only for the most promising nodes.
The judge isn’t just a cheap substitute — it’s the guidance signal. Because pairwise verdicts arrive in seconds rather than hours, the search gets intermediate feedback at every step instead of waiting for slow evaluation runs to complete before exploring further.
The numbers
The headline result: using o3-mini as the base coding agent, SIFT pushed a DGM-style initial harness from a 14.2% baseline to 35.1% on the full Polyglot benchmark within 30 expansion steps — beating DGM’s 30.7%, which required 80 nodes of search. The winning run completed in under five hours of wall-clock time, 42 CPU hours, and roughly $150 in total API credits, judge costs included.
The runs at different price points are just as telling:
- Qwen3-30B agent with a Qwen3-480B judge: 31.1% for $34.30 total ($4.10 of it judging) in 6.7 wall-clock hours / 224 CPU-hours
- Qwen3-30B agent with a gpt-5.4 judge: 32.0% for $33.70 in 6.8 hours / 188 CPU-hours
- o3-mini agent with a gpt-5.4 judge: 35.1% — roughly a tenth of DGM’s CPU budget
Two structural findings stand out beyond the raw scores. First, harnesses transfer: an agent harness discovered while self-improving with one coding model carried its gains to other models, meaning the discovered scaffolding is capturing something general rather than overfitting one backbone. Second, the exploration machinery matters — removing the exploration parameter dropped o3-mini results from 35.1% to 30.1%, confirming the gains come from genuine search, not from a lucky lineage.
Why this matters
The significance is less “35.1%” and more “$150.” Sakana AI frames this as part of its Recursive Self-Improvement (RSI) research program — reaching the inflection point where agents write, benchmark, and verify the code of their own underlying systems. SIFT is the budget version of that thesis, and the budget version is the one that replicates.
For academic labs, a sub-$200, sub-250-CPU-hour entry point turns self-improving agents from an industrial project into a course-project-sized one. For the field, it widens the population of people who can audit, reproduce, and stress-test recursive improvement loops — exactly the cohort you want examining a capability with obvious safety implications. The paper’s own safety section notes that recursive loops amplify standard agentic risks, and that unbounded self-modification is the failure mode to watch.
There are honest limitations. The judge can be wrong, and a strong-but-biased judge would bias the search; the authors note that nodes consistently dominated in both judge rank and measured accuracy could be pruned, and that future judges could be made agentic — running lightweight probes or inspecting traces rather than reading patches cold. Polyglot is also a benchmark, and benchmark-tuned self-modification is still benchmark-tuned.
But the direction is clear. The expensive part of self-improvement was never the writing of the code — it was deciding which version of yourself to keep. SIFT shows that a reasonably good, extremely cheap answer to that question is enough to keep the loop turning.