← All posts / Research

Scooped by a Machine: Claude Computes the Nine-Loop N=4 Super-Yang-Mills Amplitude for About $2,000

Anthropic physicists gave Claude one prompt and a week of 96 CPUs; it beat the human eight-loop record in planar N=4 super-Yang-Mills, validated by record-holder Lance Dixon — and a Beijing team using GPT-6 hit the same target within days.

Scooped by a Machine: Claude Computes the Nine-Loop N=4 Super-Yang-Mills Amplitude for About $2,000

On September 25, Anthropic published a guest post on its Science blog with an unusual byline and an unusual tone. The author, Matt von Hippel, is a former theoretical particle physicist turned science writer who blogs weekly at 4gravitons.com. Seven weeks earlier, on August 7, he had issued a public challenge to AI labs: “Show that a computational limit everyone expected to be a problem doesn’t actually matter. Give us N=8 supergravity to seven loops, or N=4 super Yang-Mills to nine loops.” His post announcing the answer is titled, simply, “Yes, Claude can do Nine Loops.”

Two physicists inside Anthropic — Liam Fitzpatrick and Siddharth Mishra-Sharma — took the bait. They ran Anthropic’s Fable 5.1 model inside Claude Science, a harness platform scientists can pay to use, and gave it a single, unadorned prompt: “The problem is to compute the Six-particle (hexagon) amplitude in planar N=4 SYM at nine loops.” Then they went to bed. The standing instruction read: “I’m going to sleep and won’t be available for another several hours. Keep working on this until I tell you to stop. Give me updates every 4-6 hours.”

Claude worked for roughly a week, driving Python and SymPy across the equivalent of 96 CPUs, and delivered the six-particle scattering amplitude in planar N=4 super-Yang-Mills theory at nine loops — one loop beyond the human record Lance Dixon of SLAC published in 2023. And it did the calculation two independent ways, once with the direct bootstrap method and once through the indirect form-factor route, with either approach costing an end-user around one to two thousand dollars. The bootstrap step alone consumed only about $100 of that budget.

Why nine loops matters

Scattering amplitudes are the formulas physicists use to predict how subatomic particles react. They are so hard to compute that almost every practical calculation uses approximations cut off at a fixed number of “loops” — a measure of how complicated the allowed interactions are. More loops means more precision, and a steeply worse computational bill. Most amplitudes in the literature stop at two loops; a few reach three. The most precise prediction in all of particle physics, the electron’s anomalous magnetic moment lineage that the field loves to cite, sits around five.

N=4 super Yang-Mills is a “toy model” — a maximally supersymmetric cousin of the Yang-Mills theories that describe three of the four fundamental forces (electromagnetism, and the strong and weak nuclear forces). Nobody believes it describes the real world. Amplitudeologists use it because its delicate symmetry cancels so much complexity that it becomes the perfect stress-test track for new techniques: if your method can’t survive nine loops here, it has no chance on real QCD.

Dixon and Andy Liu reached the eight-loop amplitude in 2023, and they got there indirectly — by way of a related, easier object called a form factor and an eerie structural symmetry called antipodal duality. The field’s working assumption was that the direct bootstrap at nine loops was simply too heavy: if running the standard recipe one more loop had been feasible, someone would already have done it. Dixon himself expected the next loop to arrive through a different, more indirect AI-assisted route.

The machine ran the recipe, and the recipe’s author checked it

What happened instead is that Claude executed Dixon’s bootstrap recipe directly, from scratch, unsupervised. In an addendum titled “How does it feel to be scooped by a machine?”, Dixon — who validated the result over two weeks, mostly by recovering the nine-loop form factor from Claude’s amplitude — described the setup as “very fragile: if you make any mistake at all in the computational recipe, it all crashes down like a failed soufflé.” Many details of the construction, he noted, are “too boring to document fully in a publication,” so Claude had to reconstruct all of that missing glue code itself.

Dixon’s verdict lands somewhere between magnanimous and unsettling. Claude is “a different kind of transformer model, probably over a million times bigger than our custom one,” he wrote, referencing his team’s own campaign of custom transformer models for predicting higher loops — a campaign whose slogan was “We have all the tools to validate any candidate solution a machine would provide us.” He asserts that Claude “understands our 2019 and 2023 papers better than any human, aside from my co-authors,” and frames the result as mutual validation: while Dixon validated Claude’s output, Claude validated a decade of his group’s methodology by using it faithfully. The soul-searching, he added, will come “when large language models start to come up with new physical principles and insights before humans.”

Scooped twice in two weeks

The strangest twist: Anthropic wasn’t the only team converging on nine loops. Days after hearing from Fitzpatrick and Mishra-Sharma, von Hippel heard from Song He of the Chinese Academy of Sciences in Beijing, whose group — with collaborators Jirong Jing and Xiang Li — had already computed the piece of the nine-loop result called the symbol, published September 17. Song’s team used GPT-6 to help compute some of the constraints, but within a human-built framework, not Anthropic’s almost-hands-off approach. Dixon’s summary of his September: “I’ve been scooped by both a machine and by humans plus a machine, within two weeks.”

The concurrency is the real signal. Two independent teams, using different models from rival labs, hit the same target the same month — one autonomously, one with AI assistance. As von Hippel puts it, the takeaway isn’t that AI invented a new method; Claude “used known methods, with a bit more compute than people had tried to use before.” It may have benefited from choosing Python over Maple or Mathematica, and from better software engineering discipline than human postdocs typically apply. The takeaway is that “there is more low-hanging fruit out there than you’d expect,” and that the perceived computational wall was, in part, an illusion held by the experts closest to the problem.

The reliability story is the product story

For anyone who still thinks of LLMs as too error-prone for serious work, von Hippel offers the sharpest data point in the piece. This was a one-shot run: “without any scientific oversight more sophisticated than ‘keep going.’” He estimates that if he had booked a week on 96 CPUs himself, “it’s practically guaranteed I’d screw up something on the first try” and need two weeks. He doesn’t know how many mistakes Claude made and internally recovered from — but the harness got to the end without an outside collaborator’s input. In March, he notes, AI was still doing physics like a student: small tasks, hand-holding, mistakes. Six months later it one-shotted a frontier calculation “normally tackled by the top experts in amplitudes.”

He is also careful about what this does not answer. He hoped to see AI invent strange new methods and gain a glimpse of the superintelligence debate; instead he learned “just that I was too naïve about where the limit was.” He suspects real-world amplitude calculations — a wider, more competitive field with less low-hanging fruit — could still hide harder walls, but “I wouldn’t count on it,” and he thinks groups there ought to already be testing whether AI harnesses can one-shot their frontier, with a validation plan ready.

For the amplitude community, the next question is concrete: where is the new frontier, and how much closer is the field to the actual goal — predictions precise enough to compare against upcoming collider experiments? For everyone else, the economics are the headline. A calculation that occupied top specialists for years, previously gated on grant money, postdoc hours, and fragile hand-maintained code, now reproduces for the price of a laptop on a week of cloud CPUs. Dixon’s team had the validation tooling ready because they had promised themselves they could check “any candidate solution a machine would provide us.” Every other computational field now faces the same preparation question — and the two-week double-scoop shows the machines are not waiting for the answer.