Claude Breaks a 37-Year Mathematical Record: Riemann Zeta Bound Improved From 41.6% to 67.2%
An unreleased research build of Claude, told to 'take a real stab' at the Riemann hypothesis, instead improved a decades-old lower bound on zeta zeros — coordinating 60 subagents over 31 million tokens.
When Jarred Sumner, a staff member at Anthropic (and, by his own description, not a mathematician), gave an unreleased research version of Claude an “unreasonable challenge” — take a real stab at the Riemann hypothesis — nobody expected the model to succeed. The Riemann hypothesis, first posed in 1859, is one of the most famous unsolved problems in mathematics and carries a million-dollar Millennium Prize bounty. Claude didn’t prove it. But what happened instead may matter more for the future of AI-assisted mathematics: during its failed attempt, Claude discovered a genuine improvement to a decades-old result connected to the hypothesis, raising a lower bound on the fraction of Riemann zeta zeros that satisfy the hypothesis from 41.6% to 67.2%.
Anthropic published the full account on August 10, 2026, along with Claude’s paper, a Lean formalization of the proof, an informal expert note, and complete transcripts of the process.
What the result actually says
The Riemann zeta function describes the distribution of prime numbers: each zero of the function contributes successively finer detail to the sequence of primes. The Riemann hypothesis conjectures that all of the zeros that determine the primes lie along a single vertical line in the complex plane — the so-called “critical line.”
Nobody has proven that. But mathematicians have spent decades quantifying how many zeros provably sit on the line. Before Claude’s run, the best-established lower bound was that at least 41.6% of zeros satisfy the hypothesis — a figure that had stood as the state of the art for years. Claude’s new result raises that guaranteed proportion to 67.2%.
Crucially, this is not a case of an AI reinventing known mathematics. Claude’s proof draws on a specific, recent line of research: techniques introduced by Montgomery in 1973, which originally assumed the hypothesis was true, and subsequent work by Aryan and by Baluyot, Goldston, Suriajaya, and Turnage-Butterbaugh that made those techniques work without the assumption. Claude combined these with a 2000 paper by Bombieri to surpass the previous bound.
The technical heart of the proof is striking in its boldness. Claude formed a suitable space of functions with a quadratic form induced by Weil, with positive-definite subspaces arising from zeros on the line and negative-definite subspaces from zeros off it. It then wrote down an inequality on the rank of the quadratic form using first- and second-moment information. As Anthropic’s write-up notes, the key step was “the courage to treat the entire space, with positive- and negative-definiteness taken into account together, and with the quadratic form allowed to be non-diagonal.” That willingness to leave the diagonal, partitioned territory where prior work had stayed is what let Claude push past 41.6% on the strength of existing results.
How Claude got there: 650 failed ideas and 60 subagents
The methodology section of Anthropic’s post is as remarkable as the math. The entire discovery unfolded across two Claude Code sessions consuming a total of 31 million output tokens.
In the first session, Claude generated and tested roughly 650 candidate ideas for attacking the Riemann hypothesis. None of them worked. Sumner prompted Claude to try again, and the model spent a day and a half coordinating approximately 60 subagents. Between them, the subagents ran 2,400 shell commands and wrote hundreds of Python scripts, executed thousands of numerical checks against known zeta zeros, and refereed one another’s work.
Perhaps the most human detail in the entire account: Sumner’s contribution during this phase consisted mostly of messages of encouragement — variants of “keep going” and “believe in yourself.” According to Anthropic, this appears to have actually helped. Claude was initially skeptical it could make meaningful progress on such a problem, having internalized from its training both the difficulty of open mathematical problems and the limitations of AI models. “Perhaps Claude, like many of us, underestimates the rate of AI progress,” the company writes.
The division of labor among the subagents is worth noting. Out of 60, two were responsible for developing the key mathematical ideas, 13 contributed ideas to those agents, 30 attempted but failed to develop new ideas, 13 served as validators checking correctness, and the final two helped write the initial paper.
Claude then did something no purely human research process would spontaneously do at that speed: it stress-tested its own result. Subagents reviewed the proofs, searched for counterexamples, downloaded 54 papers from the arXiv to verify the finding hadn’t already been published, and independently re-proved the result from scratch. The model volunteered to write up its findings as a paper — and recommended that a human number theorist validate them.
Human verification, and a formal proof
That validation came from two directions. Levent Alpöge and Ralph Furman, mathematicians at Anthropic, studied Claude’s paper to understand the new results and their relation to prior work. In parallel, Claude worked with staff member Eric Easley to produce a Lean formalization of the result, which passes the standard validation tool comparator — meaning the proof is machine-checkable, not merely human-reviewed.
Anthropic also brought in outside experts: Brian Conrey and Dan Goldston, two authorities in the field, examined the paper on short notice and produced an informal note stating the proof concisely. The post was updated on August 13 with a revised version of Claude’s paper that provides a clearer proof and additional historical context.
Important caveats remain. Anthropic is explicit that the techniques Claude used are not expected to lead to proving the Riemann hypothesis itself. The result is a bound improvement — a significant one, built on human mathematics, but not a crack in the hypothesis. Forbes’ coverage framed it as breaking “a math record that stood for 37 years,” which is fair as a headline, though the deeper story is about how the record fell.
Why this matters beyond number theory
Three takeaways stand out from this episode.
First, agent orchestration is becoming a legitimate research method. A single model instance chatting about math is one thing; a model that spawns 60 subagents, assigns them roles (ideators, validators, writers), has them cross-check numerical evidence against known zeros, and conducts an automated literature review of 54 arXiv papers is something qualitatively different. Anthropic frames the result as “a data point on the agent-orchestration approach to hard math problems” — and it’s a strong one.
Second, AI is now extending human mathematics rather than just reproducing it. Claude didn’t discover a new field or invent new machinery from nothing; it found a novel combination of existing results — Aryan, Baluyot et al., and Bombieri — that no one had put together, and had the “courage,” as Anthropic puts it, to work with non-diagonal quadratic forms on the full function space. Extending the impact and reach of mathematicians’ ideas is precisely the augmentation scenario AI optimists have promised.
Third, the sociological detail matters. The discovery was made by a non-mathematician prompting a research model with encouragement rather than technical guidance. All mathematical choices were left to Claude. If meaningful mathematical progress can be triggered by a layperson’s prompt and a chorus of “keep going,” the bottleneck in AI-assisted research is shifting decisively from human expertise to model capability and compute.
The Riemann hypothesis remains open, and Claude’s 67.2% bound will need to withstand sustained scrutiny from the number theory community. But the shape of this result — an unintended byproduct of a failed attempt at a harder problem, self-verified, formally proven in Lean, and validated by leading experts — offers the clearest glimpse yet of what AI-native mathematical research looks like.