Neurosurgeon + GPT-5.6 Solve Crouzeix's Conjecture, a 22-Year-Old Math Problem
A Beijing neurosurgery resident with no formal math training proved the 22-year-old Crouzeix's conjecture using a 16-hour autonomous GPT-5.6 Sol run — and an independent proof landed on arXiv eight days later.
One of the strangest and most consequential mathematics stories of the year unfolded over the past three weeks: Dr. Shanmu Jin, a postdoctoral researcher and neurosurgery resident at Peking Union Medical College Hospital, proved Crouzeix’s conjecture — an open problem in matrix analysis dating to 2004 — with the decisive help of ChatGPT running GPT-5.6 Sol in autonomous mode for roughly sixteen hours.
The story, first detailed in a widely shared SIAM News essay by Alex Townsend and Anne Greenbaum on August 14, upends assumptions about who can contribute to frontier mathematics and what role large language models can play in producing genuinely new proofs rather than polishing existing ones.
What Crouzeix’s Conjecture Says
For a square matrix $A$, the numerical range $W(A)$ is the set of all values $x^*Ax$ taken over unit vectors $x$ — a compact convex region in the complex plane that always contains the eigenvalues. For nonnormal matrices, eigenvalues alone can tell an incomplete and misleading story; the numerical range captures considerably more of the matrix’s behavior.
Crouzeix’s conjecture, posed by French mathematician Michel Crouzeix in 2004, states that for every polynomial $p$:
$$|p(A)|2 \le 2 \max{z \in W(A)} |p(z)|$$
In words: the numerical range is a “2-spectral set” — a region on which scalar function values control the matrix norm of the polynomial applied to the matrix, with a universal constant of 2. The bound is sharp, and the constant 2 cannot be improved.
The inequality matters far beyond pure curiosity. As Townsend and Greenbaum note, if a polynomial or rational function approximates a function uniformly on $W(A)$, a Crouzeix-type bound transfers that scalar approximation error into a norm bound on the matrix function $f(A) - r(A)$. This is a standard route from complex-plane approximation to the analysis of matrix functions, GMRES convergence, and polynomial and rational Krylov methods — the numerical linear algebra machinery underneath much of scientific computing.
A Gravy Train of Partial Results
The problem attracted a substantial literature. Crouzeix himself proved in 2007 that replacing the constant 2 with 11.08 always suffices. In 2017, Crouzeix and Palencia reduced the universal constant to $1 + \sqrt{2} \approx 2.414$ — tantalizingly close, but not the conjectured 2. That same year, the American Institute of Mathematics hosted a dedicated week-long workshop in San Jose on the conjecture, organized by Mark Embree, Michael Overton, and Anne Greenbaum, attacking it from theoretical, numerical, dilation-theoretic, and inverse-numerical-range angles. Despite compelling numerical evidence and two decades of effort, nobody could close the gap.
The Unlikely Prover
Dr. Jin’s route into matrix analysis began, of all places, with research on transcranial ultrasound. Teaching himself the subject, he encountered Crouzeix’s conjecture and was drawn to its unusually simple statement and the visual geometry of the numerical range. His formal mathematical education consisted of the courses ordinarily taken by science students — he studied geology as an undergraduate before earning an M.D. “Everything beyond that has been self-taught,” he told Townsend and Greenbaum by email.
For about a year before the breakthrough, Jin had been routinely asking ChatGPT to solve the conjecture. It would either return a bogus proof or stall at a missing lemma. On July 30, 2026, it linked to a preprint posted July 27 titled The Numerical Range Is a 2-Spectral Set, claiming a complete solution. Townsend and Greenbaum initially read the preprint with skepticism — but after a few hours realized “the argument was the real deal.” Both of them, along with Michel Crouzeix himself, have checked the proof thoroughly and believe the manuscript is correct.
The Sixteen-Hour Run
The key theorem emerged during an approximately sixteen-hour autonomous run of GPT-5.6 Sol in ChatGPT Work mode. The setup is as interesting as the outcome. Jin adapted a prompt OpenAI had used to solve the Cycle Double Cover conjecture. The prompt explicitly denied the system access to the public web and other external contexts. Jin started the run and did not intervene.
The prompt required a branching portfolio of genuinely different approaches. It instructed ChatGPT to spin up many subagents to individually explore proof strategy, with the explicit constraint that they not converge prematurely on the same attractive idea. Candidate proofs were repeatedly subjected to adversarial audits, and strategies were ruled out only with counterexamples. Finally, the prompt instructed the model not to give up until a complete proof survived checking.
The decisive theorem stunned the verifiers: rather than requiring the much stronger estimates everyone expected, the argument reduces the problem via a careful sampling strategy to a surprisingly simple positivity condition. Jin’s public repository includes the exact prompt, successive manuscript versions, a Lean formalization, and an axiom audit — making the work unusually open to examination. He has been strikingly modest, emphasizing there was “certainly an element of luck” in finding the key idea — all while keeping up a clinical neurosurgery schedule that sounds grueling.
An Independent Proof, Eight Days Later
On August 4 — only eight days after Jin’s first preprint appeared — Emiel Lorist and Felix Schwenninger posted an independent five-page proof on arXiv (arXiv:2608.03841). Their argument is different and strikingly short: it combines the classical double-layer potential representation with a perturbation lemma for 2-dilations, and may be friendlier to those who tried and failed to prove the conjecture in the past. Notably, Lorist and Schwenninger also disclose that ChatGPT 5.6 was used to explore proof strategies.
Jin’s reaction to the independent proof was, according to Townsend and Greenbaum, pure delight: he described it as an interesting complement to his own method and was pleased that two genuinely different proofs emerged almost simultaneously.
Why This Matters Beyond One Conjecture
The mathematics community’s own verdict is the story here. Townsend and Greenbaum — top experts in numerical linear algebra — write that it is “remarkable” a longstanding conjecture was first solved by someone with no specialized training in mathematics, working with a language model. Only a few years ago it was difficult to imagine an AI system contributing the decisive idea in a proof of a major conjecture. Now it has happened.
They expect episodes like this — outsiders making major contributions with AI assistance — to become increasingly common, even as the burden of verification continues to fall largely on experts. Their closing question is the one academia is now forced to confront: whether the system has the capacity to absorb and digest the flood of results that may be coming. As they put it, this is “surely the beginning of a golden age for results in applied mathematics” — but what it means to be an applied mathematician may be about to dramatically change.
Two caveats are worth holding onto. First, verification remains human, expert, and slow; a proof only counts once specialists like Crouzeix, Townsend, and Greenbaum have checked it. Second, the result validates a particular way of using these models — long autonomous runs, adversarial self-auditing, no web access, forced exploration of heterogeneous strategies — not casual prompting. The craft is in the orchestration, and Dr. Jin, a doctor with no math degree, may have just written one of its founding documents.
Sources
- [1] https://alextownsend.net/essays/SIAMNews_CrouzeixConjecture.pdf
- [2] https://arxiv.org/abs/2608.03841
- [3] https://www.scmp.com/tech/tech-trends/article/3363966/chinese-doctor-stuns-maths-world-cracking-decades-old-problem-using-chatgpt
- [4] https://interestingengineering.com/ai-robotics/china-doctor-cracks-math-problem-with-chatgpt
- [5] https://github.com/jinshanmu/CrouzeixConjecture