The Wall Fell in 14 Months: GPT-6 Astra Solves the Last FrontierMath Tier 4 Problem, and Epoch Declares Saturation
Epoch AI has declared FrontierMath Tier 4 fully saturated after GPT-6 Astra cracked the final holdout — a problem by combinatorialist Jay Pantone. The benchmark built to resist AI for years went from 5% to 98% in fourteen months, and the evaluation frontier is already moving to Erdős problems in Lean.
Benchmarks are supposed to die slowly. FrontierMath Tier 4 — the research-mathematics tier that Epoch AI built with more than seventy contributing mathematicians precisely so it would resist AI systems “for years” — is instead dead in fourteen months. On September 10, 2026, Epoch announced that every problem in Tier 4 has now been solved by an AI model, with OpenAI’s GPT-6 Astra taking down the last problem standing. The benchmark’s final holdout was authored by combinatorialist Jay Pantone, and the top score on the leaderboard now reads 98%.
The announcement closes one of the fastest saturation arcs in the modern benchmarking era. When Tier 4 launched on July 11, 2025, the best model of the day solved roughly 5% of it. Fourteen months later the ceiling is essentially gone, and Epoch has formally marked the suite saturated — the label it applies when a benchmark no longer meaningfully discriminates between frontier systems.
What FrontierMath Tier 4 actually is
FrontierMath is not a word-problem set. Epoch AI assembled hundreds of unpublished, extremely difficult problems across mathematics, written specifically so they could not be leaked into training data, and organized them into tiers of escalating difficulty. Tier 4 is the apex: research-level questions that push into genuine mathematical territory, with programmatically verifiable numerical answers. That verifiability is what made FrontierMath so useful as a benchmark — no human grading, no partial credit, no arguing with the referee.
The problems are hard in a way that resists brute force. Tier 4 items were designed to require the kind of multi-hour, multi-step reasoning that expert mathematicians deploy on real research questions. In early 2025, before the reasoning-model era fully arrived, the entire suite looked like it would hold for a long time. o3’s celebrated 25% on the full FrontierMath set in December 2024 was considered a shock; Tier 4 specifically sat near zero.
The fall, in three acts
The collapse happened in stages, and each one reset expectations. The first crack came in late 2025, when reasoning models with massive test-time compute — GPT-5 Pro, Gemini 2.5 Deep Think, Grok 4 Heavy — began chipping away at Tier 4. GPT-5 Pro set a record at 13%, edging out Gemini’s Deep Think by what Epoch noted was a single problem, statistically indistinguishable from noise. The “Battle Royale” era had begun.
The second act was the shortcut controversy. Epoch documented that several early Tier 4 solves — including a pair of Pantone problems — were achieved through numerical shortcuts the problem authors had not intended. Models were, in effect, reverse-engineering the answer from the structure of the question rather than doing the mathematics. Epoch responded by auditing the entire dataset: FrontierMath v2, released in June 2026, corrected 123 problems in Tiers 1–3 and 12 in Tier 4, and removed flawed items outright. Tier 4 tightened, and scores dipped accordingly.
The third act was saturation. GPT-6 Astra launched on September 3, 2026, posting 97.6% on Tier 4 v2 — one problem short of a clean sweep. This week, that final problem fell. Epoch’s announcement carries a detail that matters to mathematicians: unlike many earlier solves, no one has reported that Astra exploited an unintended shortcut on Pantone’s last problem. The wall didn’t crumble; the last stone was pulled out.
The caveats that keep this honest
Saturation headlines deserve asterisks, and this one has several.
The funding relationship. Epoch itself discloses that OpenAI funded part of FrontierMath’s development and holds exclusive access to a portion of the dataset — a relationship that has been contentious since the benchmark’s launch in November 2024 and one Epoch has worked to make more transparent. The scores are Epoch’s, but the optics are unavoidable: the lab that helped fund the benchmark is the one that saturated it.
The runners-up are close. GPT-5.6 Sol sits at 83.0% and GPT-5.6 Terra at 68.3% on the current leaderboard. Astra finished the job, but the trend line — not any single model — is the story. Tier 4 was going to fall within months regardless of who swung the final hammer.
Saturation is not omniscience. On Humanity’s Last Exam, Astra’s 57.2% actually trails its predecessor GPT-5.6’s 65.0%. On Epoch’s new FrontierMath Erdős set — 68 Erdős problems formalized in Lean, where models must produce complete machine-checkable proofs rather than final answers — Astra officially solved just 2 of 68. Repeated aggressive attempts pushed that to 5, at a reported compute cost exceeding $220,000. Single-answer research math is effectively closed; proof construction against open problems remains wide open.
Where the frontier moves now
Epoch saw this coming. Its Open Problems track collects significant unsolved research questions with computationally verifiable solutions, and FrontierMath Erdős is the flagship successor: problems so hard that nobody — human or machine — knows the answer, formalized so that any claimed solution can be machine-checked. Astra tops that leaderboard too, but the low absolute score is the more honest indicator of where research mathematics sits relative to AI systems today.
The pattern extends beyond Epoch. The evaluation field has spent 2026 pivoting toward benchmarks that sit in the single-digit-percent regime: long-horizon agentic tasks, multi-hour coding work, formal proof construction in Lean, and genuinely open problems where no ground truth exists. The lesson of Tier 4’s fourteen-month arc is that any closed-form, single-answer test — however cleverly designed — is now a depreciating asset. Verification-resistant evaluation, not question difficulty, is the new design constraint.
Why it matters
For model developers, the practical consequence is immediate: FrontierMath scores no longer differentiate frontier systems, and citing them in launch materials is edging toward marketing theater. For everyone else, the milestone is a clean data point in the debate about how fast mathematical reasoning is actually progressing. Fourteen months is the time it took to go from “the best system barely registers” to “done, formally” on a benchmark that dozens of professional mathematicians built to resist exactly this.
It is also the quiet backdrop to a noisier week. OpenAI’s announcement that Astra had resolved the Navier–Stokes existence and smoothness problem — a Millennium Prize problem — remains contested, with priority disputes and misconduct allegations still flying. Tier 4’s saturation is a reminder that the capability underlying those claims is real and independently verified, even as the norms around announcing mathematical results continue to be fought over.
The wall fell in fourteen months. The next one is built out of Lean files and open problems, and by every current measurement, it is holding — for now.