← All posts / Research

The Erdős Gold Rush: How a Dead Mathematician's Problem List Became AI's Toughest Benchmark

Over a hundred of Paul Erdős's open problems have fallen since October 2025 — to hobbyists with chatbots, Google DeepMind agents, and OpenAI's unreleased Astra model. Quanta's deep dive explains why this quirky list became the proving ground where AI learned to do real mathematics.

The Erdős Gold Rush: How a Dead Mathematician's Problem List Became AI's Toughest Benchmark

In 2023, a British mathematician named Thomas Bloom wanted a convenient place to keep track of open problems in number theory and combinatorics, so he built himself a website. He coded it with ChatGPT’s help — a novelty at the time — and launched erdosproblems.com with a couple hundred entries, “with the expectation that maybe nobody would use it.”

Three years later, Bloom’s database has become something he never imagined: the informal benchmark where the world’s most powerful AI systems compete to prove they can do genuine mathematics. Over 111 problems on the list flipped from “open” to “solved” between 2024 and August 2025 alone, and the pace has only accelerated. Google DeepMind used Gemini to systematically evaluate 700 open conjectures from the database. OpenAI’s unreleased Astra model resolved ten long-standing problems in mathematics and theoretical computer science in a single announcement. And a striking share of the results have come not from corporate labs but from hobbyists and undergraduates wielding publicly available chatbots.

The story of how this happened — told in detail in Konstantin Kakaes’s August 3 feature for Quanta Magazine — is one of the clearest windows we have into how AI is actually changing mathematical research, for better and for worse.

The man behind the list

Paul Erdős was one of the most prolific mathematicians in history. The itinerant Hungarian lived out of a suitcase, published over 1,500 papers with hundreds of collaborators, and rattled off new problems constantly — in papers, in letters, in conversation — often attaching cash bounties payable to whoever solved them first. He died in 1996, but a nonprofit foundation in Iowa has kept the bounty promise alive.

Erdős’s problems have a particular flavor. They are typically simple to state, deep to solve, and concentrated in number theory, combinatorics, and graph theory — precisely the areas where large language models have turned out to be strongest at mathematical reasoning. The problems also vary enormously in difficulty and significance, which makes the list a natural ladder for a technology whose capabilities vary just as widely.

That variance is what turned Bloom’s site into a benchmark. “This kind of single repository that anyone can access — labs realized that this could effectively be a benchmark,” Jared Duker Lichtman of Stanford told Quanta. In January 2026, a 24-researcher Google DeepMind team published a paper solving four problems and recovering nine forgotten older solutions after using Gemini to work through 700 “Open”-labeled conjectures. In May, a separate 21-person DeepMind team reported that its most capable agent “autonomously resolved 9 of 353 open Erdős problems at the per-problem cost of a few hundred dollars” — a strikingly commercial framing, pricing proof-search the way you’d price cloud compute.

Hobbyists got there first

The strangest part of the story is who was doing the solving before the labs piled in. Bloom told Quanta he was surprised that despite intense attention from OpenAI, Google DeepMind, and several startups, most early results came from hobbyists and undergraduates using public models.

Consider Wouter van Doorn, who works in customer service support and nearly finished a math master’s degree a decade ago. Spurred by how quickly LLMs were improving, he took a six-month leave in 2024 to finish his own projects while he still could: “Right now I’m still better at mathematics than an AI is, but who knows what it’ll be in a year, two years, five years?” He became the fourth-most-prolific commenter on erdosproblems.com, and in November 2025 found himself defending a proof about square-free integers in the comments against another user — who turned out to be Terence Tao. Van Doorn is now a co-author with Tao on published work. “A lot of my recent papers should be mostly credited to AI,” he told Quanta. “The ideas involved were ideas I did not come up with myself.”

Then there are Kevin Barreto and Liam Price, both in their early twenties, who met on an AI Discord server in the summer of 2025 and started throwing batches of Erdős problems at GPT-5.2 in December. They learned to prompt the model in what Barreto called a very particular way — “gaslighting it into thinking the problem is easier than it actually is,” because the model wouldn’t seriously attack a problem it believed was open. On Christmas morning, Barreto posted what they believed was the first fully autonomous LLM resolution of an Erdős problem — only for another user to point out, hours later, that Erdős himself had resolved Problem 333 in a 1977 paper nobody had checked. “As someone who’s fallen for this twice now, it’s quite gut-wrenching,” Barreto wrote.

Undeterred, they cracked Erdős Problem 728 by January 4, 2026, using GPT-5.2 Pro, then had Harmonic’s Aristotle system formally certify the proof, and a collaborator had ChatGPT write up the formalization. Price, by his own assessment, lacks the mathematical training to verify the solutions himself — his method is to iteratively ask a fresh chatbot instance to check the previous chatbot’s work until it converges. With Barreto’s help, they recruited real mathematicians to check the results, and both are now co-authors with Tao and Lichtman on a May 2026 paper resolving Erdős Problem 1196, concerning primitive sets of integers.

The frontier labs take over

Everything changed on May 20, 2026, when OpenAI announced that an internal model had found a counterexample to the unit distance problem — Erdős’s famous 1946 conjecture about how many pairs of equidistant points can be placed among n points in the plane. Mathematicians had believed the conjecture true for 80 years. The model found a construction importing tools from algebraic number theory, a distant branch of math nobody had successfully applied to the problem before.

The reception was not polite applause. “This is a really impressive piece of work. … It is definitely an intimidating construction,” wrote Jacob Tsimerman of the University of Toronto. Tim Gowers — a Fields medalist — went further: “if a human had written the paper and submitted it to the Annals of Mathematics and I had been asked for a quick opinion, I would have recommended acceptance without any hesitation. No previous AI-generated proof has come close to that.” Weeks later, human mathematicians including Bloom used related techniques to disprove a version of Erdős’s sum-product conjecture.

On August 1, OpenAI followed up with “Ten advances in mathematics and theoretical computer science” — ten results from an internal version of Astra, its next major model, including solutions to three more Erdős problems. The company released the proofs with Lean formal-verification certificates and chain-of-thought walkthroughs for each, and the paper’s author line reads simply “OpenAI.” The Verge’s verdict six days later was blunt: “The AI takeover of mathematics has begun.”

What it means

The consequences for the humans involved are already visible, and they are not the ones many predicted. Noga Alon of Princeton, who has solved dozens of Erdős problems across his career, has simply stopped: “Once AI started to solve them, there is no point anymore.” Tao has stepped away from the Erdős community to focus on other work. And in July 2026, on the very day Tsimerman received the Fields Medal, mathematics’ highest honor, he announced he was leaving academia for a job at OpenAI — one of many first-rate mathematicians now migrating to AI labs, as Alon put it, because “maybe this is where the action now is.”

The picture isn’t uniformly rosy. Bloom warned Quanta about the flood of 100- to 200-page AI-generated papers where no human has read the proof: “I got AI to generate the proof and check the proof and write the paper. But no human has read it, and no human is going to read it. It’s a huge challenge now.” The database itself reflects the churn: 565 solved, 652 still open at the time of Quanta’s reporting.

Yet the collaborators who verify carefully are publishing real mathematics faster than the field can comfortably absorb, and outsiders are participating at a scale never before possible. Van Doorn’s take is the most quotable equilibrium: “If you want to play piano, you aren’t going to hire a piano-playing machine that does it better. You will play the piano because you like playing the piano.”

Erdős, who offered cash prizes for his problems and famously referred to God as the “Supreme Fascist” who kept the best proofs to Himself, would presumably have appreciated the irony: his whimsical bounty list, maintained by one mathematician and a volunteer community, is now the scoreboard in the most consequential capability race of the decade — and the machines are clearing problems priced at a few hundred dollars each.