Benchmarks You Can't Study For: Inside Artificial Analysis Intelligence Index v4.2
Artificial Analysis rebuilt its flagship leaderboard around private test sets — 40% of the weighting — added agentic knowledge-work evals, and retired the saturated GPQA Diamond. Claude Fable 5.1 holds #1 over GPT-6 Astra on the new scale.
On September 4, Artificial Analysis shipped version 4.2 of its Intelligence Index — the composite leaderboard that has become one of the most-watched scoreboards in the AI industry. The update is not a routine refresh. It redraws what the index measures, how much of it can be gamed, and quietly reorders the race at the top: Claude Fable 5.1 retains the number-one position on the new scale, ahead of OpenAI’s GPT-6 Astra, while the freshly added agentic evaluations and a doubled weighting on private test sets change what “most intelligent model” actually means.
What changed in v4.2
The Intelligence Index has always been a weighted composite. Version 4.2 aggregates ten evaluations: AA-Briefcase, GDPval-AA v2, τ³-Banking, Terminal-Bench v2.1, SciCode, Humanity’s Last Exam, GDP.pdf, CritPt, AA-Omniscience, and AA-LCR v1.1. Two of those are new, and one famous name is gone.
The most significant addition is AA-Briefcase, a private agentic knowledge-work benchmark developed by Artificial Analysis itself. It scores models on open-ended professional deliverables, combining a rubric pass rate with Elo ratings for analytical quality and presentation — the kind of messy, multi-step work that a consultant or analyst actually does, rather than multiple-choice science questions. The second addition, GDP.pdf, is a document-reasoning suite built around a 4,592-page corpus spanning ten professional domains, scored on an all-pass criterion. It tests whether a model can hold onto and reason over an entire book-length context without losing the thread.
Meanwhile, GPQA Diamond is retired. The graduate-level science question set — for years a fixture of every frontier model’s launch slide — has saturated: top models cluster so tightly at the top that it no longer discriminates. Its removal is a small milestone for the field. When a benchmark stops separating models, it stops being information.
The deeper structural change is weighting. Private test sets — evaluations whose contents are not publicly visible and therefore cannot leak into training data — now account for 40% of the index, double their share in v4.1. That is a direct response to the contamination problem that has eroded trust in public benchmarks across the industry. You cannot study for a test you cannot read.
The leaderboard: Fable 5.1 holds, Astra closes
On the live v4.2 leaderboard, the top of the chart reads: Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) at 57, GPT-6 Astra (max) at 55, GPT-6 Astra (xhigh) at 54, and Claude Opus 5 (max) also at 54, with Opus 5 (xhigh) at 53. Anthropic holds two of the top five configurations; OpenAI’s new flagship, released September 3, sits second and third.
One caution for anyone tracking scores over time: the index was re-composed, so numbers are not comparable across versions. Under the v4.1-era methodology, the same Fable 5.1 configuration scored 66 — the highest Artificial Analysis had ever recorded at the time. On the v4.2 scale it reads 57. Both statements are true; they are different rulers.
The open-weights picture is its own race. Kimi K3 (max) leads open models at 50, followed by GLM-5.3 (max) at 49 and Qwen3.8 2.4T A95B at 47 — seventeen open-weights models now populate the 53-model index. On GDPval-AA v2, the Elo-based real-world work benchmark, Meta’s Muse Spark 1.3 (max) tops the standalone leaderboard at 1720. The gap between the best open model and the best proprietary one has compressed to single digits on this scale, which is itself a story.
The asterisk that matters: cost and system scores
The v4.2 refresh arrives days after Artificial Analysis published its first independent per-task cost accounting of Fable 5.1, and the numbers complicate the victory lap. Fable 5.1 costs roughly $3.7 per Intelligence-Index task — about 20% more than its predecessor Fable 5 (~$3.1) and roughly 1.6× Claude Opus 5’s $2.34. The driver: Fable 5.1 generates approximately 1.7× the output tokens, and output is billed at $50 per million tokens, the most expensive line in Anthropic’s pricing stack. Anthropic’s headline 75% cache-read price cut ($1.00 to $0.25 per million tokens) saves about $1.40 per task on this workload — real money, but not enough to offset the extra generation. Capability rose, and on this workload, cost rose with it.
There is a second, subtler caveat. Artificial Analysis ran Fable 5.1 with Anthropic’s default server-side fallback active, and roughly 4% of benchmark output tokens were actually served by Opus 4.8 or Opus 5, routing safety-flagged requests away from Fable 5.1 itself. The number-one score is therefore a blended result — a measurement of a model plus its serving stack, not the weights alone. As every lab ships models wrapped in routers, effort dials, and safety fallbacks, “model X is number one” is quietly becoming “model X’s product configuration is number one.” That is arguably more honest about how these systems are used in production — and it is exactly what a benchmark industry needs to be explicit about.
Why this matters
Three implications follow. First, for buyers: a composite index is a blended workload, and your bill depends on your own token distribution — cache-heavy agent loops and output-heavy generation behave very differently against the same price sheet. Second, for labs: with 40% of the weighting on private evaluations, benchmark-driven training — deliberately or through leakage — gets harder, and independent measurement becomes the only check on launch-day claims. Third, for the research community: retired saturated benchmarks and agentic, deliverable-based evaluations push the field toward testing what models do rather than what they know.
The influence is already visible. When MBZUAI’s Institute of Foundation Models launched the K2 Horizon fleet this week, its own comparison tables quoted Artificial Analysis’s GDPval-AA Elo figures alongside classic benchmarks — a sign of how quickly the index’s private, agentic evaluations have become common currency.
What to watch
Expect the v4.2 leaderboard to churn as GPT-6 Astra configurations are re-run and Meta, Google, and the open-weights cohort post numbers on the new scale. The number worth tracking is not who is #1 this week — it is whether the doubling of private-test weighting changes the ordering at all. If the public-benchmark era’s leaders hold their positions on tests they could never have seen, that is the strongest signal yet that frontier capability is real rather than remembered.
Until then, one rule of thumb survives every methodology revision: label the harness, state the workload, and never compare a score across versions of the ruler.
Sources
- [1] https://artificialanalysis.ai/evaluations/artificial-analysis-intelligence-index
- [2] https://artificialanalysis.ai/models
- [3] https://artificialanalysis.ai/methodology/intelligence-benchmarking
- [4] https://fourweekmba.com/ai-claude-fable-5-1-artificial-analysis-cost-benchmark/
- [5] https://aiweekly.co/ai-news-today