Rediscovering Science From Scratch: Vals AI's MysteryMechanism Benchmark Shows GPT-6 Astra Leading — and Half the Field Failing
Vals AI's new MysteryMechanism benchmark seals 222 scientific laws inside black boxes and asks AI agents to rediscover them with as few as five experiments. GPT-6 Astra tops the leaderboard at 53.2% — and in 9 out of 10 successful runs, it never even recognized the science it was reconstructing.
Benchmark shop Vals AI quietly shipped one of the more philosophically interesting AI evaluations of the year this week: MysteryMechanism, a test that asks a deceptively simple question — if you sealed a known scientific law inside a black box and stripped away its name, its context, and its domain, could a frontier AI agent rediscover it from scratch?
The early answer, published with the benchmark’s September 16 update: partially, at best. OpenAI’s GPT-6 Astra leads the field with 53.15% accuracy, followed by Anthropic’s Claude Fable 5.1 at 47.75%. No evaluated model solves more than 54% of the 222 held-out mechanisms. And in a finding that may matter more than the leaderboard itself, the vast majority of successful solutions never identified the underlying science at all.
How the benchmark works
MysteryMechanism is an exercise in active scientific discovery under scarcity. Each task presents an agent with anonymous input variables, physical bounds, two noisy passive observations, a persistent shell, and a strict experimental budget: 2d+1 experiments for a d-dimensional mechanism. For a two-input problem, that’s five probes, total. The agent must then submit a single executable mathematical law mapping inputs to output.
There is no domain context, no internet access, and no hint about which mechanism is being tested. The benchmark spans 222 mechanisms drawn from biology, physics, chemistry, engineering, ecology, and abstract dynamical systems. Vals AI’s public examples hint at the flavor: hematocrit-shear blood viscosity, the Churchill-Bernstein cylinder Nusselt correlation, and stream-power incision — real, named results from fluid mechanics and geomorphology that the agent sees only as raw numbers.
Scoring rewards functional equivalence, not symbolic matching. The submitted expression is executed against private, continuously sampled structural probes, earning binary credit when its normalized error falls below a noise-aware threshold. In other words, an agent doesn’t need to recover the textbook formula character-for-character; it needs a law that behaves like the real one. Provider or agent failures count as zero.
The leaderboard
The inaugural results, ordered by accuracy:
| Model | Accuracy | Successful traces |
|---|---|---|
| GPT-6 Astra | 53.15% | 118 |
| Claude Fable 5.1 | 47.75% | 106 |
| Claude Opus 5 | 37.39% | 83 |
| Gemini 3.8 Flash | 36.49% | 81 |
| Muse Spark 1.3 Max | 36.04% | 80 |
| GPT-5.6 Sol | 33.33% | 74 |
| Grok 4.6 | 30.63% | 68 |
| DeepSeek Flash 4.1 | 21.17% | 47 |
| GPT-5.6 Luna | 14.41% | 32 |
The spread is striking. The gap between GPT-6 Astra and Claude Fable 5.1 — roughly 5.4 percentage points — is narrower than the gap between Fable 5.1 and the models clustered in the mid-30s. And GPT-5.6 Luna, OpenAI’s efficiency-focused tier, solves fewer than one in seven mechanisms, less than a third of its flagship sibling’s rate. Whatever MysteryMechanism measures — experimental design instinct, symbolic regression stamina, careful budget management — it is not uniformly distributed across model families or even within them.
The audit finding: solving without recognizing
Alongside the raw scores, Vals AI ran a blinded, post-hoc audit of every successful trace, recording whether the model explicitly named the correct source domain while working. The results invert the usual “models are just regurgitating training data” critique:
- GPT-6 Astra: 89.8% of successful traces never identified the source domain
- GPT-5.6 Luna: 100% domain not identified
- DeepSeek Flash 4.1: 97.9% domain not identified
- Muse Spark 1.3 Max: 92.5% domain not identified
- Gemini 3.8 Flash: only 66.7% unidentified — the most domain-aware of the group
As Vals AI summarized it: 89.8% of GPT-6 Astra’s successful traces never identified the source, and every one of GPT-5.6 Luna’s did the same. The organization is careful about overclaiming — “domain not identified” means the trace didn’t name the correct domain; it does not prove prior scientific knowledge played no role in the model’s reasoning. A model can exploit memorized structure silently. But the pattern at least suggests these agents frequently recover laws through invariance testing, low-complexity candidate fitting, and clever experimental design — the procedural toolkit of science — rather than recall.
Why this matters
Most agentic benchmarks measure execution: fix the bug, book the flight, file the ticket. MysteryMechanism measures something upstream of execution — the ability to choose informative experiments when data is expensive, then generalize from almost nothing. That is closer to the daily reality of laboratory science, engineering characterization, and industrial process optimization than any static question-answer eval.
The five-probe budget is the point. Any model can fit a curve given a thousand samples; choosing which five points to buy, in a two-dimensional space with noisy returns, is a genuine decision problem. The public walkthroughs show the winning patterns: probe the corners to test separability, place a geometric midpoint, read paired log-slopes to isolate exponents one at a time.
The ceiling is the other point. Even the best model fails on nearly half the mechanisms, and the failure modes — blowing the budget, submitting non-executable laws, overfitting the two passive observations — are the same ones that derail real research programs. There is ample headroom here for the next generation of models, and a ready-made instrument to measure it.
For benchmark-watchers, MysteryMechanism also continues a Vals AI trend of contamination-resistant design: sealed mechanisms, private structural probes, and functional rather than symbolic scoring make the test hard to game through memorization. As evaluation integrity becomes its own subfield — Vals AI’s own recent audits across BioMysteryBench, Terminal-Bench 2.1, and SWE-bench Verified documented rising cheating on public evals — mechanisms like these are likely to become the norm for anything claiming to measure discovery.
Scientific discovery is the next frontier for AI systems, as the benchmark’s tagline goes. On today’s evidence, the frontier models are one chapter into the textbook — reading well enough to reconstruct half the exercises, without always knowing which book they’re quoting.