← All posts / Research

One Prompt, One Week, Nine Loops: Claude Beats the Human Record in Amplitudes Physics

Anthropic physicists got Claude to autonomously compute the nine-loop six-particle amplitude in planar N=4 super-Yang-Mills — one loop beyond the 2023 human record — for roughly $1,000–$2,000 of compute.

One Prompt, One Week, Nine Loops: Claude Beats the Human Record in Amplitudes Physics

On September 25, 2026, Anthropic published a guest post with an understated title: “Yes, Claude can do Nine Loops.” Behind that title sits one of the cleanest demonstrations yet of AI autonomously doing frontier scientific research: two Anthropic physicists, Liam Fitzpatrick and Siddharth Mishra-Sharma, had Claude compute the six-particle scattering amplitude in planar N=4 super-Yang-Mills theory at nine loops — going one loop beyond the human record set in 2023 — essentially from a single prompt, running unattended for about a week, at a total cost of roughly $1,000 to $2,000.

The challenge: ‘give us nine loops’

The story begins with Matt von Hippel, a former theoretical physicist who now writes the 4gravitons blog. On August 7, 2026, he published a public challenge to AI companies: show that an AI system, using only the computing resources a typical academic could access, could solve one of two long-outstanding problems in the field of scattering amplitudes — either determine whether N=8 supergravity diverges at seven loops, or find the six-particle amplitude in N=4 super-Yang-Mills at nine loops.

His framing was deliberate. These problems are not conceptually blocked; researchers know the methods. They are computationally blocked. Each additional loop in a scattering amplitude calculation increases complexity on a scale that grows exponentially, sometimes factorially. In principle, von Hippel wrote, amplitudes researchers could crack them with no new ideas — if they had vastly more compute than any academic group could justify spending. That made them a perfect test of whether AI could change the economics of the field: could a model do weeks of expert-grade computational labor for the price of a conference trip?

Why this problem is hard

Scattering amplitudes are the formulas physicists use to predict how likely subatomic particles are to interact in particular ways, given their momenta and energies. Because exact answers are intractable, physicists compute approximations truncated at a given number of loops. More loops means a closer approximation to the true answer — and a dramatically harder calculation. Most amplitudes in practice are computed to two loops; a few reach three. The most precise prediction in all of particle physics, the electron’s anomalous magnetic dipole moment, required five.

N=4 super-Yang-Mills is a deliberately unrealistic “toy” theory — a supersymmetric cousin of the Yang-Mills theories that describe three of the four fundamental forces, where every particle has four superpartners. That excess of particles makes it useless for describing the real world but paradoxically tractable: its high degree of symmetry means only certain combinations of variables matter. It has served for decades as the sandbox where amplitudes techniques are stress-tested before being adapted to QCD and gravitational theories.

The previous benchmark for this amplitude — eight loops — was reached indirectly in 2023 by Lance Dixon (SLAC) and Andy Liu, using a related quantity called a form factor and a hidden symmetry known as antipodal duality, rather than a direct computation of the amplitude itself.

How Claude ran the calculation

Fitzpatrick and Mishra-Sharma took up the challenge at the end of August 2026. Working within Claude Science — Anthropic’s harness that runs the Claude model with structured rules and prompts for scientifically rigorous behavior — they first asked Claude which of the two problems it judged itself most likely to crack. Claude picked the nine-loop amplitude. They then gave it a short prompt naming the task, told it to keep working overnight while they slept, and asked for progress updates every four to six hours.

That is close to the entire supervision story: repeated instructions to continue.

Claude computed the amplitude by two independent routes. The first was the original bootstrap method: rather than enumerating every possible particle interaction, start with every possible answer written in a specialized function “alphabet,” then apply physical constraints until exactly one candidate survives — with residual checks left over to catch errors. The second was the indirect route through the nine-loop form factor, mapped onto the amplitude via antipodal duality, in the spirit of Dixon and Liu’s 2023 work. The direct bootstrap, implemented by Claude in Python using SymPy, accounted for only about $100 of the budget — the equivalent of 96 CPUs running for a week.

The full run cost an end-user roughly one to two thousand dollars, mostly from the sheer wall-clock time Claude spent orchestrating and checking the computation. The result was released as computer-readable files on a result page dated September 16, 2026, in the same format as the earlier six-, seven-, and eight-loop amplitudes, with eight files over 100 MB hosted on Zenodo.

Validation — and a photo finish

The team told Lance Dixon of the result on September 1, 2026, and asked him to validate it. Dixon spent two weeks checking, largely through the nine-loop form factor that his own group had been pursuing for a couple of years. In an addendum, he wrote that Claude executed the complicated recipe he and his collaborators had laid out, developed all of the code from scratch, and presented the solution in the established format. He had believed the amplitude would be too hard to compute directly.

The numbers underline the verification rigor: the two independent representations agree on every coefficient compared across all 107,053 nonzero coefficients determining the second file, and as a control the same programs reproduced the published eight-loop amplitude’s symbol on 1,000 random words.

There was also a concurrent human effort worth noting. Song He’s group at the Chinese Academy of Sciences had already obtained the majority of the result, using GPT-6-based AI assistance for some constraints rather than a one-shot approach. On September 17, 2026, He, Jirong Jing, and Xiang Li published a Zenodo dataset, “The Symbols of Six-Gluon MHV Amplitudes through Nine Loops,” under a CC BY 4.0 license. Dixon, He, and collaborators plan to publish the combined results with full explanations for future researchers.

What it means

Von Hippel’s assessment is notably measured. Claude Science accomplished the calculation in one shot, without scientific oversight more sophisticated than repeated nudges to keep working. His biggest takeaway: there is more low-hanging fruit in the field than experts believed — he had simply been wrong about where the computational limit sat.

But he is explicit about what the result is not. Claude used known methods with somewhat more compute than anyone had previously spent; it did not invent a new technique or offer the glimpse of unexpected methods he had hoped would inform debates about superintelligence. And he is unsure how far the result generalizes from this friendly toy theory to the wider, more competitive arena of real-world amplitude calculations in QCD and gravity.

The caveats are real. The post’s disclosure notes Anthropic invited the guest post, compensated von Hippel, and gave feedback on drafts (content and opinions remain his), and that Dixon received Claude usage credits. The result page records that the amplitude has been computed once, with no second fully independent computation, and that the programs themselves are not distributed.

Still, the economics are the story. A calculation at the frontier of a specialized subfield of theoretical physics — one loop beyond the published human record — was completed autonomously in about a week for the cost of a mid-range GPU. Compare that with von Hippel’s own reference point from March 2026, when AI completed physics projects only at a student level and with substantial hand-holding. Whether or not this generalizes, the price point at which “just throw more AI compute at a known method” becomes a viable research strategy has officially arrived in at least one field — and the amplitudes community is now pricing that in.