← All posts / Tools

One Engineer, $120,000 in Tokens, 832,378 Lines of Rust: Inside GitHub's Agent-Led Copilot Runtime Rewrite

GitHub let Copilot's own agents rewrite the Copilot agent runtime from TypeScript to Rust — in place, on main, shipping continuously for 14.5 weeks. The full numbers, the regressions, and what it says about agent-led engineering.

One Engineer, $120,000 in Tokens, 832,378 Lines of Rust: Inside GitHub's Agent-Led Copilot Runtime Rewrite

The most consequential rewrite in developer tools this year wasn’t done by a platform team over two quarters. It was done, in large part, by the very product being rewritten. In a engineering post-mortem published on the GitHub Blog, Microsoft Distinguished Engineer Stephen Toub detailed how GitHub ported the Copilot agent runtime — the engine that executes Copilot’s coding agents — from TypeScript/Node.js to 100% Rust, with AI agents writing most of the code, one human operating the control loop, and production shipping continuously the entire time.

By August 21, 2026, the runtime stood at 832,378 lines of production Rust, backed by 468,689 lines of Rust unit tests and 174,675 lines of end-to-end TypeScript tests. The rough bill: about $120,000 in attributed token spend plus roughly three weeks of one developer’s time across a fourteen-and-a-half-week porting window. As Toub put it: “Agents moved the price to where the project became tenable.” A rewrite of this scale in a live system would previously have demanded a full team for a year or two — and, he admits, it probably should have lost the resource argument.

A moving target: 130,000 lines became 430,000

The initial May 2026 estimate put the runtime at roughly 130,000 lines of TypeScript. That number turned out to be “reasonably accurate, but wildly misleading.” While the port was underway, code kept cascading downward from the TUI layer, entire components were ruled in-scope late, and tens of developers kept merging hundreds of pull requests per week — constantly adding new TypeScript. In the end, Toub estimates approximately 430,000 lines of production TypeScript passed through the port, roughly 3.3× the original scoping figure.

The churn didn’t stop there. During the port the runtime took in ~300,000 new production TypeScript lines and shed ~430,000, while ~1,200,000 Rust lines entered and ~365,000 left. To outside observers, TypeScript volume looked stable for most of the effort — masking the fact that porting was silently keeping pace with incoming feature work.

In-place, atomic, always shippable

GitHub rejected the big-bang rewrite in favor of in-place atomic replacement: each pull request ported one component, replaced the TypeScript with a thin shim calling into Rust via napi-rs, and deleted the old implementation in the same atomic change. The main branch never stopped shipping — 135 releases went out during the window (100 pre-releases, 35 stable), averaging about 1.3 releases per day alongside roughly 1.3 port pull requests per day.

The strategy produced an unexpected safety property: because ported TypeScript was deleted on arrival, any rebase that touched already-ported code generated a guaranteed merge conflict — one branch editing it, the other deleting it. Incoming changes to ported code could not slip through unnoticed.

The hardest single artifact was session.ts, an organically grown ~30,000-line backbone threading through nearly every subsystem. The porting session that took it on spent its first 56 minutes doing nothing but reading — 122 tool calls before creating anything — then split the file into slices delegated to 15 child sessions, each with its own worktree, branch, and agent loop, spawned in seven waves over about three hours. Ten of the fifteen ran on GPT-5.6 Sol, five on Claude Opus 4.8. The full run lasted 25 hours.

What 12.7 million session events reveal

The post’s most unusual contribution is its audit of the agent session logs: 12,760,995 events, 31,247 user-role messages, 1,385,214 assistant messages, and 1,130,921 tool calls — 61% of which came from subagents rather than the main thread. Of those user messages, Toub personally typed or spoke only about 2,600, one in twelve. His role, as he describes it, was “less assign-a-task-and-wait and more operate the control loop”: inspecting results, challenging technical decisions, enforcing quality gates, and pushing when agents treated an intermediate stopping point as the finish line.

The economics hinge on prompt caching. The port achieved a 96.22% prompt-cache hit rate, with cache writes at 3.07% and genuinely fresh input at just 0.71%. With providers commonly discounting cached reads around 90%, that cache discipline is, in Toub’s words, “the reason the economics of long autonomous sessions hold together at all.” Context was auto-compacted 5,116 times; one sessions-infrastructure pull request alone compacted 647 times over its multi-day lifespan.

How much unsafe, and how many regressions

Skeptics of agent-written Rust assume the borrow checker gets shrugged off at scale. It didn’t. The entire runtime crate contains 158 unsafe blocks across only 36 files, and every one sits at a genuine interop boundary: C ABI entry points (51), Windows API calls (49), POSIX/libc (46), SQLite’s C API (7), dlopen (4), and one process-environment mutation. Notably, none of the known regressions involved an unsafe block, and unsafe is absent from the model clients, MCP layer, agent layer, and prompt layer entirely.

There were regressions — dozens traced by September 14, all fixed, mostly correctness bugs with a smaller set of performance issues. The lessons section reads like a handbook for agent-led engineering: state the goal completely (“port component X” was read by agents as hot-paths-only until the end-state was made explicit); end-to-end tests are non-negotiable because they’re the oracle; protect the oracle from the agent — the thing changing the implementation must not be able to silently weaken the tests that define correctness; translate first, redesign second; and when a failure mode appears twice, promote it into standing instructions, an eval, or the harness itself.

Why it matters beyond GitHub

The payoff is architectural: the runtime now exposes a C ABI and loads in-process in all six Copilot SDK languages (C#, TypeScript, Python, Rust, Go, Java) — no Node.js, no V8, no second process to supervise, which Toub calls the single most common friction partners reported. A runtime instance costs a fraction of what it did, so hosts can run far more concurrent sessions, and the binary can now target everything from cloud to embedded devices.

The deeper signal is economic. A project category — wholesale rewrites of live, business-critical systems — that was previously unfundable became a one-engineer, five-figure-token operation. Toub is careful not to overclaim (“I didn’t simply ask an AI agent to ‘port this whole codebase’”), but the conclusion is hard to escape: the constraint on large rewrites is shifting from writing the code to specifying, validating, and vouching for it. “After three months of watching agents rewrite the very engine that runs them,” Toub writes, “I’m looking forward to seeing just how far it goes.”