← All posts / Research

The Last AI Built by Humans: 35 Chinese Researchers Publish a 75-Page Roadmap for Recursive Self-Improvement

A 35-author team spanning Shanghai Jiao Tong University, Tsinghua, Shanghai AI Lab and Theseus Labs has published the first autonomy-centered roadmap for recursive self-improvement — five levels from executing prescribed fixes to AI that rewrites its own improvement process — and a new benchmark metric showing where frontier models actually stand.

The Last AI Built by Humans: 35 Chinese Researchers Publish a 75-Page Roadmap for Recursive Self-Improvement

The most consequential AI may not be the one that solves the hardest problems — it may be the one that solves the problem of making better AI. That is the wager behind “The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement”, a roughly 75-page position paper and survey published on arXiv (2609.11873) by a 35-author team led from Shanghai Jiao Tong University together with Tsinghua University, Shanghai AI Laboratory, Theseus Labs, ModelBest, Tencent Hunyuan and other academic and industrial partners. First posted September 10 and revised September 15, 2026, it has quickly become one of the most-discussed papers of the month on Hugging Face and alphaXiv.

Its central question is blunt: everyone from OpenAI’s Noam Brown to Demis Hassabis now says recursive self-improvement (RSI) — AI that materially accelerates the development of AI — is the frontier’s top priority. But what does that actually mean, how would we know when we have it, and where do today’s systems really sit? The paper’s answer is a five-level autonomy ladder, a new capability-maturity metric, and a sober warning: almost everything marketed as “self-improving” today is level-two at best.

Why “The Last AI Built by Humans”?

The title is deliberately provocative. If improvement becomes genuinely recursive — each generation of AI making the next generation smarter, faster, and cheaper to build — then at some point humans stop being the principal engineers and become something more like regulators, customers, or gardeners. The paper does not claim that point has arrived. Instead it tries to do what no widely-adopted framework has done: give the field a shared vocabulary for measuring progress toward it, so that “our system self-improves” stops being marketing and starts being a falsifiable claim.

The corresponding author is Xuanhe Zhou (Shanghai Jiao Tong University), with Yi Duan, Ying Liu and Zirui Tang among the equal-contribution lead authors, and a team drawn from academia and industry — including practitioners from Tencent’s Hunyuan group and ModelBest contributing deployed-system case studies. The work is anchored by a public project page (theseus-labs-rsi.github.io) and positions itself explicitly against hand-wavy RSI discourse: every level in its roadmap comes with an evaluation protocol for deciding whether a system has reached it.

The Headroom-Closed Index: A Ruler for Uneven Progress

Before defining RSI, the paper measures the terrain. Its authors introduce the Headroom-Closed Index (HCI) — a normalized score of how much of a benchmark family’s available “headroom” (distance from the entry-year frontier to a perfect score) the best models have closed. The formula rescales model scores against the 90th-percentile model in the benchmark’s first year in the dataset, then aggregates families per domain with square-root weighting by coverage.

The findings are the paper’s empirical hook. By 2026, advanced mathematics and graduate-level science reach HCI values of 86.4 and 85.8 respectively, while broad knowledge sits at 77.2 — but other domains lag badly, and the timing of gains differs sharply across capability areas. The implication: foundation-model progress is radically uneven. Static benchmark leaderboards flatter models precisely in the domains where headroom has been exhausted, while interactive, stateful, tool-using, long-horizon capabilities — the ones RSI actually needs — remain far from closed. This unevenness, the authors argue, is exactly why “improvement” must become a persistent, autonomous process rather than a series of one-off training runs.

The paper also grounds its motivation in current industrial reality: Kimi K3 activates only 16 of 896 experts (claiming ~2.5× scaling efficiency over K2), Qwen3.8-Max activates 95 billion of 2.4 trillion parameters, and OpenAI reports GPT-5.6 Sol designing and running hundreds of experiments on its own speculative-decoding draft model. Development burden, not capability, is the bottleneck — which is the case for handing improvement itself to machines.

The Five Levels of RSI Autonomy

The core contribution is the autonomy ladder. It takes the improvement loop — change something, measure it, keep what works — as the unit of analysis, and asks, at each level: which decisions inside that loop has the system taken over from humans?

  • B0 — In-Task AI Improvement. The baseline. The model improves within a single task instance (chain-of-thought refinement, self-correction on one prompt) but nothing persists. Today’s reasoning models live here.
  • L1 — Autonomy over Improvement Execution. The system executes prescribed improvements: data cleaning, hyperparameter tuning, kernel optimization, evaluation runs, deployment tuning. Humans decide what to improve; the machine handles the how at the execution level. The paper catalogs working industrial systems at this level across data, training methods, platforms, evaluation-safety, and deployment optimization.
  • L2 — Autonomy over Improvement Strategies. The system chooses how to improve: searching over prompts, agent harnesses, even model and training configurations. Systems like Promptbreeder, GEPA and MPO are placed here — autonomously mutating and selecting improvement strategies, but against a fixed, externally-specified fitness metric. Crucially, the paper finds today’s most-hyped “self-improving agents” remain at L2 because the objective and acceptance rules are still set externally.
  • L3 — Autonomy over Future Learning Experience. The system generates its own training experience — adaptive task generation, self-play, autonomous practice through environment interaction. It decides not just how to improve but what to practice next.
  • L4 — Autonomy in Deployment and Environmental Adaptation. Deployed systems distill trajectories, revise their own agent architecture, and selectively retain and ship updates — learning where they actually operate.
  • L5 — Recursive Meta-Improvement. The system improves the improvement process itself: rewriting its own search procedures, improving how successors are evaluated, even revising research goals and policies. This is the level at which “the last AI built by humans” language becomes literal.

The ladder’s discipline is that each level is defined by decision rights transferred, not by benchmark scores — making claims comparable across labs and preventing the familiar trick of rebranding in-context learning (B0) as self-improvement.

Four Domains, Four Different Clocks

A separate chapter applies the framework to four application regimes, showing that RSI arrives at very different speeds depending on feedback availability, verification cost, and deployment constraints:

  • Science — hypothesis modules, experimental agents and reflection systems that can be evolved, but where verification is slow and expensive.
  • Embodied intelligence — evolving curricula, skills, world models; simulation makes experience cheap but transfer to reality is the choke point.
  • Software engineering — the most RSI-ready domain, with executable tests providing cheap, trustworthy verification; evolving coding-agent harnesses is already industrial practice.
  • Healthcare — the most constrained: safety, privacy and regulatory overhead mean improvement loops stay human-gated for the foreseeable future.

The industry chapter grounds all of this with named case studies: Theseus Labs’ environment–data–model co-evolution, Lark’s enterprise data foundations for reliable RSI, Humanlaya’s delivery-driven data quality loops, ModelBest’s zero-human industrial pipelines, Tencent Hunyuan’s experience-driven self-improvement, and an agent-native research lab built on verifiable research infrastructure.

Eight Open Problems — and a Sober Verdict

The paper’s final chapter resists triumphalism. It identifies eight research directions, among them: cross-component diagnosis (an observed failure rarely tells you which component should change — data, context, tool interface, model, or evaluator, a problem the Humanlaya case makes concrete); scalable verification under autonomous practice; managing compounding interactions when multiple components adapt simultaneously; and long-horizon evaluation to test whether “improvements” transfer beyond the setting that produced them.

The verdict embedded in the framework: genuine recursive self-improvement has not been achieved anywhere. Even the strongest demonstrated systems are automating execution and strategy (L1–L2) while humans retain objective-setting, evaluation and acceptance. The bottlenecks motivating the survey — resource-intensive foundation-model development, dependence of scalable learning on useful experience and reliable verification, and the continuing human cost of diagnosing and shipping updates — “remain only partially resolved.”

Why It Matters

The paper lands in a week when RSI discourse has gone mainstream: OpenAI’s leadership calling AI-improves-AI its top research priority “by a wide margin,” Senate briefings on capability trajectories, and frontier labs racing to automate their own research pipelines. Against that noise, a rigorous taxonomy from a leading Chinese academic-industrial consortium does two useful things. It gives regulators and safety researchers a shared scale for asking how much autonomy has actually been delegated — the question that matters for oversight. And it gives engineering teams a checklist discipline: before claiming self-improvement, specify which decisions in your improvement loop the system actually makes, and show the retention evidence.

The title’s premise — that humans might one day build their last AI from scratch — remains speculative. But the roadmap to test it is now on paper, five levels deep, with evaluation protocols attached. Everything AI has achieved so far, the authors write, is but a drop in the ocean; the question their ladder makes answerable is who, or what, fills the rest.