Neurosurgeon With No Math Degree Cracks 22-Year-Old Crouzeix's Conjecture With a 16-Hour ChatGPT Run
Beijing neurosurgery resident Jin Shanmu set GPT-5.6 loose on Crouzeix's conjecture — a numerical linear algebra problem open since 2004 — and a 16-hour autonomous run returned a proof that Michel Crouzeix himself has verified.
On July 30, 2026, Alex Townsend — a professor of applied mathematics at the University of Washington — opened arXiv the way he did most mornings. For the past year he had kept a peculiar habit: every so often, he would feed Crouzeix’s conjecture to GPT-5.6 and see whether the model could make progress on one of numerical linear algebra’s most stubborn open problems. In hundreds of attempts, the answers came back riddled with logical gaps or dead-ended at the same key lemma.
That morning was different. The conjecture had a proof on arXiv — and the author was not a mathematician. It was Shanmu Jin, a postdoctoral researcher and resident neurosurgeon at Peking Union Medical College Hospital in Beijing, with no formal training in advanced mathematics. His only collaborator: GPT-5.6, which had worked the problem autonomously for roughly 16 hours, without human intervention, inside ChatGPT’s work mode.
By mid-August the story had traveled from Chinese tech media to SCMP to mathematicians’ social feeds, and the community’s reaction — in Townsend’s and Anne Greenbaum’s words — was simply “shocking.” Perhaps most striking of all: among those who checked the manuscript and confirmed it was correct was Michel Crouzeix himself, the French mathematician who posed the conjecture 22 years ago.
An outsider’s path into matrix analysis
Jin’s résumé has nothing to do with mathematics. His undergraduate degree is in geology. He later switched to medicine, earned an MD, and landed in neurosurgery. “All other mathematical knowledge is self-taught,” he wrote matter-of-factly in an email to Townsend and Greenbaum, two of the field’s leading experts.
His route into the conjecture came from an unlikely source: transcranial ultrasound research. Trying to model how ultrasound propagates through the complex structure of the human skull, Jin stumbled into matrix analysis — and onto the Crouzeix conjecture, a problem whose statement is remarkably simple but whose proof resisted the field’s best efforts for two decades. The combination of elementary statement and deep difficulty was exactly what hooked him.
Why Crouzeix’s conjecture matters
To understand the fuss, it helps to know what the conjecture says. Matrices are the working language of applied mathematics — from quantum state evolution to Google’s PageRank to the giant parameter iterations behind modern LLM training. But non-normal matrices behave erratically; taming them is a persistent challenge.
In 2004, Michel Crouzeix conjectured that for any polynomial p on the complex plane and any matrix A, the spectral norm of p(A) can be bounded by twice the maximum value of p on W(A), the numerical range of A. Formally:
‖p(A)‖ ≤ 2 · maxz ∈ W(A) |p(z)|
with 2 being the optimal constant. If true, the inequality lets you convert scalar approximation errors on the complex plane directly into norm bounds on matrix functions — a tool of decisive importance for analyzing matrix functions, GMRES-style iterative methods, and Krylov subspace techniques.
Progress was grinding. Crouzeix himself proved the bound with constant 11.08 in 2007. In 2017, the American Institute of Mathematics convened a week-long workshop in San Jose, after which Crouzeix and Palencia reduced the universal constant to about 2.414 (1 + √2). And there it stalled — the door to the constant “2” stayed closed for another decade, until this summer.
The prompt engineering behind the proof
The most practically interesting part of the story is how Jin got GPT-5.6 across the finish line. He didn’t treat the model as a question-answering machine. He turned it into a virtual mathematics research institute, adapting the prompt OpenAI famously used when its models attacked the Cycle Double Cover conjecture. His setup had four strategic pillars:
- No internet. The prompt explicitly cut the system off from the public web and external context. Jin didn’t want the model parroting the field’s history of failed approaches — he wanted original thinking from the axioms up.
- Divergent sub-agents. He instructed ChatGPT to spin up many parallel lines of attack, with an explicit warning against premature convergence on the first attractive idea.
- Adversarial auditing. Candidate proof strategies were repeatedly attacked by other agent branches; a line of inquiry died only when a counterexample to its approach was found.
- Refusal to quit. The prompt ordered the model not to give up before producing a complete proof that survives extreme logical stress-testing.
Then Jin pressed Enter, went back to his clinical work, and didn’t intervene. GPT-5.6 ground away for about 16 hours and came back with a proof. Crouzeix reviewed the manuscript and confirmed it was correct — although, importantly, the work has not yet passed formal journal peer review.
A second, independent proof appeared the same week
The story has one more twist. Around the same time Jin’s result was circulating, arXiv also received “A solution to Crouzeix’s conjecture” (arXiv:2608.03841) by E. Lorist, which combines tools developed for the weaker estimates with a new perturbation lemma. Two independent claims on a 22-year-old problem, landing within days of each other — one from a career mathematician, one from a doctor with an LLM — is the kind of coincidence that makes you wonder what else is quietly becoming tractable.
Both proofs still await the slow machinery of formal peer review, and healthy skepticism is warranted: claimed proofs of famous conjectures have collapsed before under scrutiny. But the early verdict from the conjecture’s own author counts for a great deal.
What this means for AI-assisted mathematics
Strip away the novelty of the protagonist and the pattern is the important part. The winning formula wasn’t raw model capability alone — it was:
- A motivated domain expert who understood the problem’s landscape well enough to know what a serious attempt looked like
- Careful orchestration: isolation from stale prior work, forced exploration of multiple proof strategies, adversarial self-review, and persistence
- Long autonomous runtime — 16 hours of uninterrupted compute, a scale of “effort” no human mathematician can match
This is the second time this year a frontier model has taken down a named open problem, after OpenAI’s disproof of an 80-year-old conjecture in discrete geometry in May. The trajectory is clear: models are graduating from proof assistants and hypothesis generators to primary engines of mathematical discovery — provided someone knows how to aim them.
Jin’s result also reframes the “who gets to do mathematics” question. The barrier to entry for frontier research is shifting from credentials and years of specialized training toward problem taste, orchestration skill, and the judgment to verify what comes back. A neurosurgeon with a geology degree and a well-structured prompt just beat a decade of workshop-organized expert effort. The tools were the great equalizer; the insight to use them this way was the differentiator.
Whether Crouzeix’s conjecture is finally settled will be decided by peer review over the coming months. But the meta-result — that a self-taught outsider with a scalpel and a subscription can now seriously attack the hardest problems in a field — already stands.
Sources
- [1] https://eu.36kr.com/en/p/3937408047053960
- [2] https://www.scmp.com/tech/tech-trends/article/3363966/chinese-doctor-stuns-maths-world-cracking-decades-old-problem-using-chatgpt
- [3] https://interestingengineering.com/ai-robotics/china-doctor-cracks-math-problem-with-chatgpt
- [4] https://arxiv.org/abs/2608.03841