Edited After Publication: Inside the Quietly Shifting Benchmarks of OpenAI's GPT-6 Astra Launch
A Fortune investigation using Internet Archive snapshots shows OpenAI changed at least six GPT-6 Astra benchmark numbers after the launch blog went live — halving Astra's hallucination rate, dropping Anthropic's math score by 10 points, and re-running the post's deployment twice before anyone could read it.
On September 3, 2026, OpenAI published the blog post announcing GPT-6 Astra — its 100,000-GPU flagship and the model CEO Sam Altman says marks the start of the “AGI era.” But according to a Fortune investigation published September 4 and updated September 5, the announcement the public eventually read was not the one OpenAI first published. Using Internet Archive snapshots, Fortune reporter Emily Forlini documented that OpenAI changed at least six evaluation metrics on the page after it went live — in every contentious case, in a direction that made Astra look better and its closest competitors look worse.
The result is the most detailed public case study yet of a practice researchers call “benchmaxxing”: quietly re-running evaluations under shifting conditions until the numbers tell the story a launch needs. And it arrives at a moment when benchmark scores have become the primary currency of competition between frontier labs.
A launch post that would not stay published
The rollout itself was strange. OpenAI planned to publish the Astra announcement at 2 p.m. ET. The post did go up shortly after 2 p.m. — and was then retracted. When OpenAI’s X account tweeted the link at 3:32 p.m., it returned an error. At 3:50 p.m., Sam Altman posted it himself, writing, “We hit a little snag getting the blog post deployed, but it is really great.” Many readers still couldn’t load it. About an hour later, the page finally became widely visible.
OpenAI told Fortune the retraction was for reasons it could not disclose, insisting they were unrelated to benchmark figures. (The company first blamed a content-management-system bug, then an internet outage.) But the archival record shows what changed between the version published at 2 p.m. and the one that stabilized hours later — the metrics.
What actually changed
The changes Fortune documented, snapshot by snapshot:
Hallucination rate. In the first archived snapshot at 2:23 p.m., Astra’s hallucination rate was 4.2%. It stayed there through five snapshots. In the sixth snapshot at 5:20 p.m. — after the page became widely viewable — it had been halved to 2%. The score for OpenAI’s own previous model, GPT-5.6 Sol, simultaneously worsened, dropping from 12.2% to 9.4% (a lower hallucination rate is better, so Astra’s improvement and Sol’s worsening both flattered the new model). As of Fortune’s publication, both numbers had been changed back to their original values: 4.2% and 12.2%.
FrontierMath Tier 4 (v2). Astra’s own math score of 97.6% never changed. But the comparison scores did. Anthropic’s Claude Fable 5.1 dropped from 87.8% in the 2:23 p.m. snapshot to 78% by 5:17 p.m. — nearly ten percentage points — before settling back at 83%. GPT-5.6 Sol went from 83% to 80.5% and back to 83%. For a window of hours, the published page made Astra’s math lead look dramatically larger than it does today.
ExploitBench. On OpenAI’s internal cybersecurity evaluation, GPT-5.6 Sol’s score jumped from 5.5% to 11.5% in later versions. OpenAI told Fortune it is investigating reverting that number, saying the 11.5% result reflects a reasoning level not commercially available for Sol.
Coding. Astra’s coding score got a marginal bump, from 57.7% to 57.9% — negligible, but OpenAI cared enough to swap in the bigger number.
ARC-AGI-3. Even the pre-publication embargoed draft Fortune received listed Astra at 98.6% on ARC-AGI-3. The live blog now says 99.99%. The Arc Prize Foundation’s own independent assessment put Astra at 99.9% — but only when given a “particularly powerful harness,” the bundle of tools a model uses to complete tasks. Under the benchmark’s standard harness, it scored 63%.
Not every edit flattered OpenAI. Two Anthropic scores on the healthcare-focused HealthBench Professional actually improved across versions: Fable 5.1 rose from 56.6% to 58.1%, and Opus 5 from 54.5% to 56.4%.
OpenAI’s response
A company spokesperson told Fortune: “We care deeply about getting evaluations right. Most evaluations have noise within a few percentage points based on the exact checkpoint, scaffold, and evaluation run used in reporting. For our launch blog, we made fixes to ensure the numbers represent our best estimate of available model performance, so that users can make meaningful comparisons.” On the draft-to-final changes, the spokesperson said adjustments “are normal” because evals are always verified before publication.
There is real technical truth here. Different research teams at OpenAI own different metrics and report them to a central team for publication. Scores genuinely vary with checkpoint, reasoning level, harness, and grading configuration — and rival-model numbers are typically pulled from published leaderboards, not re-run in-house. The blog itself carries a disclaimer: “Evaluation scores are the maximum at any effort.”
But that is precisely the problem the episode exposes. If multiple numbers can all be defended as “accurate” depending on conditions, then the choice of which accurate number to publish — and when to change it, and in which direction — is not a neutral technical decision.
The benchmaxxing question
Stanford researchers Anka Reuel and Mike Hardy, of the Stanford Intelligent Systems Laboratory and Stanford Trustworthy AI Lab, told Fortune the pattern fits a known industry practice: maximizing scores by re-running evaluations under different conditions. “This can be done in a very tight timeframe, and it’s better for their marketing,” they said. They also noted that the GPT-6 Astra system card — the document meant to provide technical depth on how evaluations were performed — offers “barely any details about the evaluation” for the internal hallucination benchmark, not even the number of test items.
Vincent Sunn Chen, who leads benchmark and evaluation research at Snorkel AI, was more sympathetic to OpenAI, noting that scores often shift in final launch logistics because checkpoint, configuration, harness, and grading setups are “typically still shifting in the final days before a launch.” But he argued the industry needs a norm: when a company revises published benchmark numbers, it should report what changed, so researchers can interpret the results.
There is precedent for worse. In 2025, Meta denied boosting Llama 4’s scores by benchmarking an internal version rather than the released one — until former chief AI scientist Yann LeCun admitted the company had “fudged” the results.
Why it matters
Benchmark numbers are not decoration in this market. They decide enterprise procurement, shape hiring narratives, and anchor the valuation story ahead of OpenAI’s possible 2027 IPO. When a lab can halve a hallucination rate, shave ten points off a rival’s math score, and revert both — all silently, on a live launch page — the leaderboards become marketing collateral with extra steps.
None of this means GPT-6 Astra is not a genuinely strong model. Independent evaluations, including Artificial Analysis’s coding-agent index and CodeRabbit’s code-review study, suggest it is. It means the measurement layer of the AI industry is more editable than the marketing implies — and the only archive keeping score is the Internet Archive’s.
The cheap fix is the one Sunn Chen proposed: a changelog for benchmark numbers. Until labs adopt one, every launch-day score should be read the way you’d read any number that quietly changed after publication — as a claim under negotiation.