25x in Six Months: Anthropic's Own CI Collapse Shows the Hidden Bill for Agentic Coding
Anthropic engineers now ship 8x more code per quarter and Claude authors 80% of it — but the blog's honest postmortem reveals CI jobs exploding 25x, three failing patches, and a redesign that took one engineer three weeks instead of a quarter.
Every vendor demo of agentic coding shows the same highlight reel: an agent spins up, writes a feature, passes its tests, and the pull request merges in minutes. On September 14, 2026, Anthropic published something rarer — the receipt. In an engineering postmortem titled “Agentic coding is straining CI,” staff engineer Sachin Malhotra walked through what happens to the unglamorous plumbing of a software organization when code generation stops being the bottleneck. The numbers are startling even by the standards of a company that builds the agents.
Anthropic engineers on average now ship eight times as much code per quarter as they did during the 2021–2025 period, and Claude authors roughly 80% of it — a figure the company first disclosed in May and reconfirmed here. Claude also “plays a large role in reviewing and approving PRs.” The volume of tests in the codebase grew 10x over the same window, while the number of engineers grew only nominally. The result: continuous-integration job volume increased 25x in six months.
The part nobody budgets for
The punchline of agentic coding is that writing code is no longer the constraint. The fine print is that everything downstream of writing — code review, merge queues, and above all CI — inherits the full multiplied load. Anthropic runs a deterministic test impact analysis service that decides which tests run on which pull request, based on past results and package relevance. It is the piece of infrastructure that makes “don’t run every test on every PR” possible, and it nearly became the next single point of failure.
The service’s original design was a singleton: a “listener” process recording results from every CI run, and a “selector” reading that history to pick tests for newly opened PRs. Because per-test history needed a single writer, the service could not be horizontally sharded. When multiple CI jobs started arriving every second, the listener began falling behind the PR queue — and in an AI-native SDLC, lag compounds brutally. Twenty minutes of listener lag can mean tens of thousands of test-result updates not being applied to the selector, which then makes stale decisions about what to run: flaky tests block merges, regressions slip through gaps, and on-call engineers get paged at 2 a.m.
Three patches, 70 days, 29 days, less than a day
What follows is one of the most candid scaling narratives a frontier lab has published. By October 2025 the service was already paging its team two days straight. The fixes came in a familiar descending arc:
- Patch 1 — a bigger machine. They doubled the cores. Everyone knew it was fleeting; it bought roughly 70 days.
- Patch 2 — sharding. The insight was that the listener didn’t need one writer for all results, only one writer per package. An internal long-running Claude Tag session — which had been monitoring the service for months and pinging its owner whenever backlog exceeded 50,000 jobs — generated the code to split each package’s state into its own shard. It bought 29 days.
- Patch 3 — daily restarts. By March the process was hitting its memory limit by mid-afternoon most weekdays. Restarts bought less than a day, and worse, each restart left the service slightly further behind.
The detail worth pausing on: the persistent recommendation to stop patching and overhaul the service came from Claude itself, in the monitoring sessions. The humans “usually settled on another patch” — a very human failure mode that agentic tooling is supposed to correct, and eventually did.
The redesign
The eventual fix took the agent’s advice: the singleton got an in-memory data store, with listener workers appending results to a journal statelessly, a small consumer rolling the journal into per-test history every few seconds, and the selector reading from the store. Any worker can now process any result, which makes the whole path horizontally scalable. The distributed version is more expensive to run — but it is stable, profileable, and no longer a lottery.
The postscript is the most quietly radical line in the piece: the redesign took a single engineer three weeks; a year earlier, the same project would have taken a quarter. The tool that strained the infrastructure also absorbed the cost of replacing it. Even the fine-tuning of journal sizing and worker counts “Claude did largely autonomously.”
Why this matters beyond Anthropic
Two structural shifts are hiding in this postmortem, and both apply to any team adopting coding agents seriously.
First, agents change the shape of pull requests. Claude prefers smaller, more granular PRs — which is good hygiene but multiplies the number of CI jobs per unit of delivered work. Meanwhile the activity floor rises because agents push code overnight and on weekends, while bursts still track human working hours, since humans drive and approve a meaningful share of PRs. CI load becomes both higher and less predictable.
Second, the economics of “over-engineering” have inverted. Malhotra’s explicit advice: whether you build or buy, assume your architecture will see 25x load within two quarters, and start accounting for 10–20x perceived scale in v0 designs “as long as your budget allows for it.” Keep state out of processes from day one, instrument services so agents can act as “eyes and ears” and hill-climb fixes autonomously, and never run a critical service as a single instance you can’t measure or canary.
The honest ledger
Anthropic has every incentive to sell the acceleration and less incentive to publish the induced costs, which is exactly why this post is valuable. The 80%-Claude-authored figure gets the headlines, but the transferable lesson lives in the boring middle of the stack: verification infrastructure must now scale as a first-class product of agentic coding, not an afterthought. Test impact analysis — a niche discipline most organizations have never needed — is on its way to becoming standard equipment, because the alternative is paying compute and wall-clock time to run every test against every one of an ever-growing stream of machine-written PRs.
“Always plan for the exponential” is the author’s hard-won conclusion. For engineering leaders mapping out 2027 budgets, the Anthropic postmortem offers a rare, quantified preview: the agents will write your code, and then they will write the infrastructure that survives your code. The organizations that thrive will be the ones that noticed the second part early.