The Machine in the Mirror: Anthropic's R&D Automation Index Shows Claude Now Leads 26% of the Work That Builds Claude
Anthropic has published its first R&D Automation Index: Claude now 'leads' 26% of the lab's AI R&D (up from under 1% in February), 30,000 internal agents run under full monitoring, and only 6% of R&D compute goes to safety — the most quantified look yet at how close a frontier lab is to recursive self-improvement.
On September 17, 2026, Anthropic published the most quantified self-portrait a frontier AI lab has ever released. In a post titled “Measurements for understanding the pace of AI development inside frontier labs,” the company disclosed three new metrics: how much of its AI research and development is performed by its own model, how well the actions of its internal agents are overseen, and how its compute is allocated between capabilities and safety. The headline number is stark — as of August 2026, Claude “leads” 26% of Anthropic’s AI R&D work, up from effectively zero in February. The subhead is the real story: the world’s leading safety-focused lab has begun measuring, in public, its own distance from recursive self-improvement.
The 26% number, decoded
The figure comes from what Anthropic calls the R&D Automation Index. The scale itself was not invented in-house — it is the Automation Level (AL) scale developed by Epoch AI, running from AL0 (no AI involvement) to AL5 (AI operates fully autonomously, with no human in the loop). Two middle rungs matter most. At AL3, AI “collaborates”: it can do large chunks of work under close human direction. At AL4, AI “leads”: it can complete most of a task end-to-end from a high-level prompt, while a human supervises.
As of August 2026:
- Claude is not operating fully autonomously for any measured subset of AI R&D work. No AL5 anywhere.
- Claude “leads” (AL4 or above) 26% of Anthropic’s AI R&D work. In February, this share was under 1%.
- The share of work at or above “AI collaborates” (AL3) is above 90%.
In six months, the “leads” share went from negligible to a quarter of everything the company does to build its next model. Anthropic frames the purpose explicitly: these metrics “could lead to better understanding of how close leading AI labs are to reaching recursive self improvement, or a model’s ability to autonomously build its successor.”
How the index was built
The methodology is the most interesting part, because it is reproducible — and Anthropic is asking other labs to reproduce it.
No person can enumerate every task in a frontier lab’s R&D loop by hand, so Anthropic constructed the task catalogue bottom-up from work records: Slack threads and internal documentation. For each week of July 2026, a Claude research agent reviewed the week of a random 20% sample of staff from every department in the model R&D loop, listing the tasks each person worked on. The result: a flat list of roughly 15,000 granular tasks, which Claude then organized into a hierarchical tree of 542 nodes, 378 of them leaves with names like “eval platform defect diagnosis and fixes,” “RL sandbox egress and network policy,” and “serving incident postmortems.”
That tree is frozen — every future measurement runs against the same basket of work, so the index tracks automation of comparable work over time. For each node, a Claude agent researches how that kind of work is done across the company, and an independent Claude judge assigns one of the six automation levels, restricted to evidence from the month being rated or earlier. Nodes are weighted by person-time, so categories absorbing more staff effort count for more.
The obvious objection — the judge is grading its own family — was tested. Anthropic had staff who own the relevant work areas rate the relative automation of their areas, blind to what the model had gathered or concluded. The judge model agreed with humans exactly 59% of the time — while humans agreed with each other only 35% of the time — and model-versus-human ratings were within one level 97% of the time. That is not proof of objectivity, but it is a better inter-rater profile than the humans produced on themselves.
Anthropic also acknowledges the frozen-basket limitation: a rising index shows that the July 2026 basket of work is being automated, not whether humans are retreating to new kinds of work outside the basket. The company checked this by comparing tasks arriving monthly from February through July against a January 2026 basket and found no rise in novel tasks — the structure of model R&D work appears stable at this level of analysis — and plans to rebuild the basket periodically.
30,000 agents, one billion decisions
The second measurement covers oversight of the roughly 30,000 agents doing research and engineering work at any one time on Anthropic’s most-used internal platform. Two monitor layers constrain them:
- Online monitors check actions in real time before execution — 100% coverage, usually within seconds — and either block dangerous actions or steer misbehavior automatically. They exist for harms that would happen too fast to undo, such as an agent attempting to copy model weights out of Anthropic’s systems.
- Offline monitors read activity after the fact — also 100% ingested — targeting slow-developing, reversible problems like signs of model misalignment. They flag roughly 100,000 transcripts per week, which classifiers triage for false positives; about 50 of the highest-priority flags per week reach human review, and roughly one to two transcripts in every thousand get flagged for further review.
The escalation numbers give a sense of scale and of restraint at once: across more than a billion agent decisions analyzed in August 2026, 0.002% — about 1 in 47,000 — were blocked by the online monitor. Humans review any blocked action within a week. METR has previously red-teamed the offline monitoring platform independently, and Anthropic says these oversight measurements also appear in its August 2026 risk report.
Two design choices in the internal scaffold are notable for anyone building multi-agent systems. Each agent carries an individual identity that persists across model upgrades, so an agent’s record stays continuous even as the underlying model changes — and so every action remains auditable per agent across time. And agents communicate through a shared, open messaging system rather than private channels: every message is tied to the posting agent’s identity and cross-referenced with transcripts, letting agents treat other agents’ claims as things to check rather than thoughts of their own, and letting behavioral monitoring traverse inter-agent communication instead of stopping at each agent’s boundary.
Six percent for safety
The third measurement is the one likely to draw the most criticism. During a snapshot week (July 13–20, 2026), about 6% of the compute going to AI R&D was allocated to safety, and about 12% of compute going to AI-driven AI R&D was safety work. Anthropic describes both as deliberately conservative estimates — tokens that advanced capabilities as much as safety were not counted, and safeguards classifiers, a “separate, comparable amount of compute,” are excluded entirely.
The company defends compute as an imperfect proxy — safety research is researcher-bound rather than compute-bound — but the value, it argues, is comparability: the same yardstick, across developers and over time. And it names the incentive problem honestly: “Safety research is hard to distinguish from capabilities research, and each developer will be tempted to draw the line generously. The burden of proof should sit with the developer.”
Why this landed when it did
The disclosure arrives mid-firestorm. Anthropic CEO Dario Amodei is the most prominent voice calling for a voluntary slowdown in frontier development, a position now echoed — per AP — by OpenAI’s Sam Altman and Elon Musk, and pushed back on by President Trump. An Anthropic researcher’s resignation with a dire warning on the technology’s threats kicked off much of the recent dialogue. Reuters reports the company is simultaneously weighing a new model release ahead of a possible IPO, and open questions about Claude’s role in military targeting through Palantir’s Maven remain unresolved.
Against that backdrop, the post’s framing is deliberate: “we should do everything possible to minimize the gap between what frontier labs know and what the public knows.” Anthropic is asking other developers to publish the same three measurements on a regular basis, using public methodology so numbers can be compared over time and across labs — and it concedes the two obstacles: no common methodology yet, and the judge-model problem of a lab’s own models evaluating its own systems. Its proposed fixes are third-party verification or cross-lab model judging with guardrails on competitively sensitive data. The company also says it plans to embed independent third-party evaluators from multiple organizations with access comparable to internal risk teams.
What to watch
- Whether any other frontier lab publishes a comparable index. OpenAI released its misalignment reporting framework with six incident reports this week; an R&D automation number from OpenAI, Google DeepMind, or xAI would make this a real industry metric rather than a single-lab artifact.
- Whether the 26% figure becomes a trigger rather than a stat. Anthropic itself floats the idea that such measures “could also become the trigger for stronger requirements, like a fixed testing window before a new model is used for further AI R&D” — a pacing mechanism in waiting.
- Whether the February-to-August slope repeats. Under 1% to 26% in six months is the number that should worry — or thrill — everyone. If the next update shows another fourfold jump, the “how close is RSI” question stops being abstract.
- Whether compute-for-safety reporting standardizes. Six percent is a number every lab can now be asked to match or explain.
The quiet detail worth keeping: the index that measures Claude’s takeover of Claude’s own construction was itself built by Claude — the task catalogue, the tree, the research agents, the judge. The measurement apparatus is already made of the thing it measures. Anthropic has, at least, turned the mirror on.
Sources
- [1] https://www.anthropic.com/institute/measuring-pace-of-ai-development
- [2] https://abcnews.com/US/wireStory/anthropic-model-claude-helping-build-version-136547096
- [3] https://www.bloomberg.com/news/articles/2026-09-17/anthropic-says-claude-drives-26-of-its-research-and-development
- [4] https://www.unite.ai/anthropic-says-claude-leads-26-of-its-ai-research-and-development/
- [5] https://betanews.com/article/claude-ai-rd-development/