Claude Pushes Riemann Zeta Bound From 41.6% to 67.2% — the First AI Breakthrough Past 50%
An unreleased Claude model improved a century-old lower bound on Riemann zeta zeros from 41.6% to 67.2% using 60 subagents and 31 million tokens — the largest single-step advance in a generation.
On August 10, 2026, Anthropic published a research note that sent ripples through both the AI and mathematics communities. An unreleased research version of Claude had improved a longstanding lower bound on the fraction of zeros of the Riemann zeta function that lie on the critical line — pushing it from 41.6% to 67.2%. The result drew over five million views within hours of announcement.
The Riemann hypothesis, open since 1859, is one of the seven Millennium Prize Problems and carries a one-million-dollar bounty from the Clay Mathematics Institute. It concerns the distribution of prime numbers: the zeta function’s zeros encode increasingly fine structure in the sequence of primes, and the hypothesis asserts that all non-trivial zeros sit on a single vertical line in the complex plane, known as the critical line. Nobody has proven that all zeros lie there, and nobody has disproven it either.
What Claude did was not prove the hypothesis. It improved a related bound — the proportion of zeros that mathematicians can rigorously certify as lying on that critical line. And that proportion, after more than a century of incremental human effort, just took its largest single jump in history.
A 112-Year Chain, Broken Open
The 41.6% figure that Claude surpassed was not a recent result. It was the endpoint of a 112-year chain of incremental advances, each representing a serious paper in analytic number theory:
| Year | Mathematician(s) | Proven Lower Bound |
|---|---|---|
| 1914 | Hardy | Infinitely many zeros on the line (no proportion) |
| 1942 | Selberg | A positive but unspecified proportion |
| 1974 | Levinson | More than 1/3 (≈33.3%) — the mollifier method |
| 1989 | Conrey | More than 2/5 (40%) |
| 2011 | Bui, Conrey, Young | ≈41.05% |
| 2020 | Pratt, Robles, Zaharescu, Zeindler | More than 5/12 (≈41.67%) |
| 2026 | Claude (unreleased) | 67.2% |
Read the last two rows together. From Conrey’s 40% in 1989 to the Pratt–Robles–Zaharescu–Zeindler result of 5/12 in 2020, the world’s analytic number theorists collectively moved the constant about 1.7 percentage points over 31 years — roughly five hundredths of a point per year, each step a full research paper. Claude added 25.6 percentage points in roughly a day and a half.
That is not a marginally faster human. It is a different regime of search.
How It Happened: A Jogger’s Prompt
The experiment began casually. Jarred Sumner, an Anthropic staff member who is not a mathematician, was jogging eight days before the announcement when he asked Claude to “take a real stab” at the Riemann hypothesis itself. The model obliged. Over the next 36 hours, working through two Claude Code sessions, it consumed approximately 31 million output tokens and coordinated roughly 60 subagents that executed over 2,400 shell commands and hundreds of Python scripts.
The first session was a wash. Claude generated and tested 650 different approaches to the problem. All 650 failed. A human — Sumner — told it to try again. His contribution throughout the process, according to Anthropic’s own account, was “mostly limited to sending Claude messages of encouragement,” mostly variants of “keep going” or “believe in yourself.”
That detail has drawn considerable attention. Anthropic reports that the encouragement “seems to have helped Claude overcome some initial skepticism that it could make meaningful progress.” The model had learned from training that open problems are hard and that AI models have limitations, and it applied that prior to itself — effectively terminating its own search early. Encouragement, in this framing, was not motivation but a correction of a learned stopping heuristic. A footnote notes the same encouragement pattern was used for the Fable 5 Jacobian conjecture counterexample in July 2026. Twice is a pattern.
The Second Session: Where It Clicked
The second session was where the breakthrough materialized. Anthropic’s breakdown of the 60-subagent fleet reveals a division of labor:
| Role | Count | Function |
|---|---|---|
| Core idea developers | 2 | Produced the key mathematical insights |
| Idea contributors | 13 | Fed approaches to the core two |
| Failed explorers | 30 | Attempted new ideas, none landed |
| Validators | 13 | Checked correctness, refereed each other |
| Writers | 2 | Drafted the initial paper |
Half the fleet — 30 of 60 agents — produced nothing usable. Two agents out of sixty generated the actual result. That is not an inefficiency to be engineered away; it is what search looks like at the frontier of mathematics. Any orchestration design that optimizes for every agent contributing is optimizing for the wrong objective.
The Mathematics: Old Pieces, New Assembly
Claude did not invent new mathematics. It combined two existing lines of work that no human had assembled together. The first was a series of papers by Baluyot, Goldston, Suriajaya, and Turnage-Butterbaugh that made Montgomery’s 1973 techniques work without assuming the Riemann hypothesis — a crucial prerequisite, since Montgomery’s original methods assumed the hypothesis was true, rendering them useless for proving anything about it. The second was a 2000 paper by Bombieri.
The technical core: Claude constructed a space of functions with a quadratic form induced by Weil, where positive-definite subspaces arise from zeros on the critical line and negative-definite subspaces from zeros off it. It then wrote down an inequality on the rank of that quadratic form using first- and second-moment information.
Anthropic credits the genuine insight not as a new technique or computational brute force, but as what it calls a temperament: “The courage to treat the entire space, with positive- and negative-definiteness taken into account together, and with the quadratic form allowed to be non-diagonal, is in some sense the step that allows Claude to achieve the conclusion.” Handling the non-diagonal case is unpleasant rather than impossible — and machines have an advantage in unpleasantness tolerance.
Verification: Stronger Than Most, Not Complete
The verification process was more rigorous than most AI-mathematics claims, though not equivalent to formal peer review. Two Anthropic mathematicians — Levent Alpöge and Ralph Furman — studied and validated the paper, writing an informal note stating the proof concisely for experts. Brian Conrey, who set the 40% record in 1989, and Dan Goldston, who co-authored the unconditional machinery Claude’s argument builds on, examined the paper on short notice. The people best positioned to spot an error were the ones whose own work was being extended.
Claude also produced a formally verifiable Lean proof that passes the standard validation tool, comparator. Subagents independently ran thousands of numerical checks against known zeta zeros, searched for counterexamples, downloaded 54 papers from arXiv to confirm the finding wasn’t already published, and re-proved the result from scratch before escalating to humans for final review.
What was not done: no journal submission, no formal peer review, and — critically — the generating model is unreleased and unnamed, so nobody outside Anthropic can reproduce the run. The Lean file is public and checkable, but the experiment that produced it is not reproducible in the scientific sense. The proof is more auditable than most human papers; the capability claim is less falsifiable than almost any of them.
The Cost of the Leap
Anthropic put a concrete number on the computational spend: 31 million output tokens across two sessions. At Claude Sonnet 5’s published rate of $10 per million output tokens, that is roughly $310 in output billing alone — before input tokens, cached context, or the overhead of 2,400 shell commands. At Opus-class rates, the figure is several times higher. Real API spend for a comparable run likely lands in the low-to-mid four figures.
Set against a result that moves a constant number theorists have pushed on since Hardy’s 1914 theorem, that is a remarkable trade. Set against the 650 failed ideas that preceded it, it is also a reminder that the token budget went overwhelmingly into things that didn’t work — a ratio anyone budgeting agentic research runs should plan for.
A Year of AI Mathematics
This result does not stand alone. In roughly four months, AI models have produced a string of mathematical advances: an Erdős planar unit-distance problem resolved with OpenAI’s models in May 2026, a Jacobian conjecture counterexample from Fable 5 in July (announced by the same Levent Alpöge who validated this result), Grok 4.5 on the Graffiti conjecture, and now Claude’s Riemann zeta bound.
The pattern across all of them is consistent. Models are not replacing mathematicians. They are extending the reach of existing human results by finding combinations nobody assembled. Claude didn’t invent Weil’s quadratic form or Bombieri’s paper or the Baluyot–Goldston unconditional machinery. It read all of them and noticed they fit together.
Anthropic is direct about the limitations: “We don’t expect that the techniques Claude used will lead to proving the Riemann hypothesis.” Proving 67.2% and proving 100% are different problems, and the gap is not a matter of grinding out more of the same. But for one weekend in August, an AI model pushed further into one of mathematics’ oldest open problems than any human had managed in a lifetime — and it did so because a jogger told it to believe in itself.
Sources
- [1] https://www-cdn.anthropic.com/564f962e60643842f5fcb4a17c9dbc8f608f1c37.pdf
- [2] https://techcrunch.com/2026/08/11/an-unreleased-anthropic-model-made-progress-on-one-of-maths-biggest-unsolved-problems/
- [3] https://explainx.ai/blog/claude-riemann-zeta-lower-bound-67-percent-august-2026
- [4] https://www.datacamp.com/tutorial/claude-and-the-riemann-hypothesis
- [5] https://www.neowin.net/news/unreleased-claude-model-makes-breakthrough-on-century-old-riemann-hypothesis-math-problem/