The Practitioner's Verdict: Parakhin Calls GPT-6 Max the Best Math Model Ever as GPT-6 Pro Surfaces in ChatGPT
Former Microsoft AI search chief Mikhail Parakhin says GPT-6 Max beats everything in math/ML while Fable 5.1 rules agentic work — as GPT-6 Pro quietly appears in the ChatGPT web interface.
Four days after GPT-6 Astra shipped, the frontier-model conversation has moved from launch benchmarks to something rarer and arguably more useful: a heavyweight practitioner’s honest side-by-side verdict. Mikhail Parakhin — the former CEO of Advertising and Web Services at Microsoft, the engineer who once ran Bing and built its AI search stack — posted a detailed comparison of GPT-6 and Anthropic’s Claude Fable 5.1 on September 6 that has since drawn more than 95,000 views. And buried in the same 24 hours was a quieter signal: a “GPT-6 Pro” tier began surfacing in the ChatGPT web interface for some accounts, ahead of any formal announcement page.
Both threads say something about where the frontier actually is right now — not on a benchmark chart, but in the daily workflow of someone who ships AI products.
What Parakhin actually said
The post is short but unusually dense with judgment calls. Parakhin wrote that after using GPT-6 and Fable 5.1 “extensively,” his conclusion is that “GPT-6 Max is the best model ever in math/ML, beating 5.2 Pro.” That is a remarkable claim on its face — not that GPT-6 edges out the previous generation, but that the Max configuration of the new family surpasses every model he has used on mathematics and machine-learning work specifically.
The second observation is the one practitioners will recognize instantly: GPT-6 is “extremely hampered by its tiny context space, it’s like talking to the super-smart Dory.” The Finding Nemo joke lands because it is precise. A model that is brilliant per-token but cannot hold a long working session in memory behaves like the forgetful fish — every conversation turn re-establishes context that a competitor retains natively.
His third call: Fable 5.1 remains “the best on long agentic tasks, still better at following instructions, finding elegant coding solutions.” And then the punchline that has been quoted across AI circles since: “the roles have switched: I use GPT-6 to develop a solution, then 5.1 to polish and run experiments.”
The role reversal, decoded
That last line deserves unpacking, because it inverts the division of labor that held through most of the GPT-4 and Claude 3 eras. For roughly two years the folk wisdom was the reverse: use Claude-family models to draft and design, then hand off to a GPT model for execution polish. Parakhin — who has every incentive to be neutral, having left Microsoft and having no obvious stake in either OpenAI’s or Anthropic’s fortunes — now runs the opposite pipeline. GPT-6’s raw reasoning generates the solution. Fable 5.1’s instruction-following and code taste refine it and run the experiments.
The replies under his post reinforce the pattern rather than dispute it. One developer noted you can raise GPT-6’s context window in the Codex config, with the caveat that it consumes usage linearly. Another suggested enabling the experimental context mode that substitutes notes and context search for compaction. A third described having to be “very explicit about what is allowed” for GPT-6 to persist through longer autonomous tasks, echoing a widely shared observation that GPT-6 and Astra stop earlier than users want on long-running jobs.
None of this contradicts OpenAI’s launch claims for Astra — 97.6% on FrontierMath Tier 4 v2 versus 87.8% for both Claude Fable 5.1 and its predecessor tier, per the launch coverage. Parakhin’s point is subtler: benchmark ceilings and daily usefulness are different axes, and on the second axis a model’s weakest dimension (context) can dominate its strongest (raw math).
Meanwhile: “GPT-6 Pro” appears in the wild
The second thread of the story started the same morning. Reports and screenshots described a “GPT-6 Pro” label appearing inside the ChatGPT web interface on September 6-7 — with no accompanying OpenAI blog post, changelog entry, or pricing page. Parakhin himself flagged it in his post: “This morning 6 Pro has appeared in Web interface - can’t wait to test!”
A fact-check published by explainx.ai on September 7 walks the story back to what is actually verifiable, and the honest answer is: not much, yet. No official OpenAI source confirms the tier. The sighting traces to user-spotted UI elements, the classic signature of a staged rollout, an A/B test, or an accidental feature-flag flip — any of which can vanish within days. GPT-6 Astra’s own launch on September 3 was chaotic in the opposite direction, with press coverage going live before OpenAI’s announcement page was reachable, so an unannounced tier appearing post-launch fits the messy pattern.
What is officially documented is the surrounding quota structure, which The Decoder assembled from OpenAI’s pricing pages: the $200 Pro plan includes 200 GPT-6 Pro messages per week, the $100 Pro plan gets 50 per week, Business Premium gets 50 per week, and Business Standard is capped at 15 per month. The tier is explicitly not included with ChatGPT Plus. Standard GPT-6 Astra offers roughly half the messages per five-hour window that GPT-5.6 Sol did across all plans — Plus users get an estimated 5 to 45 messages per window with Astra versus 10 to 100 with Sol.
The “best for math” problem
Explainx’s fact-check also makes a point worth keeping in mind when the next “best model for X” claim circulates. The math-leadership contest is unusually contested territory right now. In the past few weeks alone: Claude produced a 13-million-line machine-verified Lean 4 formalization of Fermat’s Last Theorem; Anthropic pushed a Riemann zeta lower bound from 41.6% to 67.2% using 60 parallel subagents; OpenAI’s Astra preview claimed ten new math advances with Lean certificates; and a viral claim that GPT-5.6 Sol helped set a prime-gap world record collapsed under scrutiny.
Against that backdrop, one practitioner’s opinion — even a practitioner with Parakhin’s engineering pedigree — is a data point about how the model performs in experienced hands, not a benchmark result. The explainer’s own guidance is the right frame: a named model, a stated harness, and a reproducible score carry weight; a screenshot and a quote do not.
Why this matters
Two things are true at once. First, the GPT-6 vs Fable 5.1 race is close enough that the winner depends on the task shape — short hard reasoning versus long agentic persistence — and a leading practitioner now openly chains the two models together rather than picking one. That composability is arguably the real headline: the frontier is becoming a pipeline, not a champion.
Second, the GPT-6 Pro sighting — whether it survives to an official launch or not — shows how model releases have become rolling events rather than discrete announcements. Tiers leak through UI flags, quotas appear in help-center pages before the launch post goes up, and the community’s verification toolkit (model-picker checks, API model lists, pricing docs) has to run in real time. For anyone building on these models, the practical advice from the availability guides is sound: wait for an official model ID, published pricing, and rate limits before writing production code against an unannounced tier.
The next data point to watch is whether OpenAI confirms GPT-6 Pro with benchmarks and pricing — and whether Parakhin, having gotten his hands on the tier he said he “can’t wait to test,” revises his verdict that GPT-6 Max is the ceiling.
Sources
- [1] https://x.com/MParakhin/status/2096606202360926224
- [2] https://explainx.ai/blog/gpt-6-pro-max-chatgpt-sighting-parakhin-math-2026
- [3] https://the-decoder.com/openai-rolls-out-gpt-6-astra-to-top-tier-chatgpt-plans-at-half-the-rate-of-gpt-5-6-sol/
- [4] https://www.atlascloud.ai/blog/tips/gpt-6-availability