Blind Benchmark Finds Frontier AI Can't Reconstruct Research Ideas: 3–15% Match Rate
A contamination-proof benchmark called Reconstruction shows seven frontier LLMs recover a paper's core idea from its bibliography alone just 3–15% of the time — while a multi-agent Swiss tournament reaches 42%.
A new blind benchmark posted to arXiv this week delivers one of the cleanest measurements yet of how far frontier language models remain from autonomous scientific discovery. The benchmark, called Reconstruction, asks a deceptively simple question: given only the reference list a paper cited before it was published, can a model reconstruct the paper’s core research idea?
For seven frontier large language models, the answer was a resounding no. Match rates against held-out ground-truth ideas clustered between 3% and 15% — a near-uniform failure that the authors argue is structural, not an artifact of prompting or model size.
What Reconstruction Tests
The task sounds tractable until you consider what it removes. A bibliography gestures toward a problem space, but the creative leap connecting existing literature to a genuinely new hypothesis is precisely what the paper itself supplies. Reconstruction is designed so that this leap — and not retrieval from training data — is the only viable path to a correct answer.
The input for each of the 643 evaluated papers across six scientific domains is deliberately stripped down: the pre-publication reference list, and nothing else. No full text. No author names. No post-publication signals.
The anti-contamination architecture operates on three levels:
- Temporal citation cutoff — no reference postdates the paper’s own knowledge horizon, preventing the model from reasoning forward in time toward work that effectively reveals the answer.
- Anonymous reference IDs — author names and identifying metadata are stripped, closing off the shortcut of recognizing a research group’s trajectory and guessing accordingly.
- Frozen per-paper bibliographies — each evaluation instance is locked so no information bleeds in at inference time.
Scoring is handled by an independent LLM judge that matches model-proposed hypotheses against the held-out ground-truth idea, allowing the benchmark to assess hundreds of papers without the prohibitive cost of expert human annotation at every data point.
Why the Uniform Failure Matters
The 3–15% range across seven frontier models is notable not for its low absolute values but for its uniformity. No single model distinguished itself dramatically at the top. The best performers edged toward fifteen percent; the rest clustered below it.
That near-uniformity rules out explanations centered on prompting strategy or model scale. The gap appears to be structural — whatever cognitive capacity produces a fifteen-percent match rate is hitting its ceiling across the board.
The result lands in the middle of a live debate about what frontier LLMs actually do when they appear to generate research ideas. If success required genuine abductive reasoning — forming a prediction from incomplete evidence — scores should vary meaningfully with model capability. The narrow spread suggests current architectures are not reliably performing that reasoning at all.
The practical implication is direct. A significant industry of “AI scientist” products now markets LLMs as hypothesis-generation engines. Many of those claims rest on evaluations that allow the model to access full paper text, author information, or post-publication signals — exactly the information types Reconstruction excludes. Scores achieved under those richer conditions cannot be assumed to reflect the model’s capacity for the inferential work itself.
The Multi-Agent Twist: Swiss Tournament Selection
The paper’s more consequential result comes from its “top 4” multi-agent pipeline. Rather than relying on a single model’s output, the pipeline:
- Solicits hypotheses from multiple models
- Subjects them to cross-model review
- Eliminates weaker candidates through a Swiss tournament — a bracket format in which hypotheses are iteratively compared until the strongest survivors emerge
No external web search is used; the pipeline draws only on what the reference list and model weights contain. The outcome: 23–42% match rates across all six domains, roughly a 2.4× lift over the best single-model baseline.
The tournament structure is doing something architecturally specific. In competition settings, individual errors cancel across rounds as stronger performers accumulate wins; applied to hypothesis generation, the format creates a computational approximation of peer review — the same distributed error-correction mechanism human scientific communities use to filter ideas before publication.
The benchmark’s most durable finding may therefore be organizational rather than parametric: the right design for AI-assisted discovery is not a bigger single model but a multi-agent structure that mirrors the adversarial validation structure of science itself. The approach has precedent — Google DeepMind’s AI Co-Scientist, released in 2025 for biomedical hypothesis refinement, applied tournament-based generation, reflection, and ranking. Reconstruction is the first benchmark to demonstrate the same lift across six scientific domains simultaneously, and under the strict no-external-information constraint that makes the result meaningful.
The 42% ceiling remains uncomfortable if the goal is autonomous scientific reasoning. Reaching it requires multiple model calls, structured tournament overhead, and coordination infrastructure that would be difficult to deploy routinely — and the computational cost of the improvement is not reported in detail in the current preprint.
How It Compares to Prior AI-Science Benchmarks
The scientific benchmark space has grown substantially since 2024:
| Benchmark | Input | Key limitation |
|---|---|---|
| ResearchBench (2025) | Inspiration papers | Rich input allows retrieval |
| HypoBench | Observed data | Synthetic ground truth; best methods recover ~38.8% |
| OpenAI GeneBench-Pro | Messy genomics data | Domain-specific |
| LAB-Bench 2 | 1,900+ biology tasks | Capability-level, not ideation |
| Reconstruction | Bibliography only | Retrieval path eliminated by construction |
What the earlier benchmarks share is input richness: the model receives enough information that partial reliance on training-data retrieval could produce meaningful scores. Reconstruction’s design eliminates the retrieval path entirely. If the paper that produced a bibliography was not in the model’s training data — and for papers published after the training cutoff, it cannot have been — then the model must reason from what prior work implies rather than from what the target paper says.
The Takeaway
The timing of the paper is pointed. Multiple frontier labs are currently publishing results suggesting AI systems can contribute meaningfully to mathematical proofs, cryptographic analysis, and literature synthesis. Those results are real. Reconstruction adds a specific, manipulation-resistant data point to the other side of the ledger: a model that impresses as a research assistant — synthesizing existing knowledge with apparent fluency — may perform far less poorly when asked to do what scientists most prize: forming a genuinely novel hypothesis from available evidence, without the answer already written somewhere it can find.
The paper is a preprint and has not yet completed formal peer review; the authors describe the current submission as a timestamped record, with full per-model breakdowns and per-domain scores expected in subsequent revisions. But the headline is already clear: the inferential step that defines scientific discovery remains, for now, a human specialty — and multi-agent peer-review-style architectures look more promising than raw model scale for closing the gap.