A 27B 'AI Scientist' From London Beats GPT-5.5 and Claude at Replicating Research
DeepMind-alumni startup Inherent released Faraday, a 27B-parameter agent post-trained with rubric-based RL that outperforms Claude Opus 4.8 and GPT-5.5 at reproducing scientific papers — by directing frontier coding agents instead of competing with them.
Of all the startups launched by Google DeepMind alumni, London-based Inherent has gotten relatively little attention. That changed on August 22, 2026, when the lab — weeks out of stealth with a $50 million seed round — published results showing that its newly released AI agent, Faraday, outperforms much larger frontier models from Anthropic and OpenAI at a task that sits at the heart of science: independently reproducing the findings of published research papers.
The result is striking not just for the leaderboard, but for the architecture behind it. Faraday is a 27-billion-parameter model — tiny by frontier standards — running on Qwen 3.6. Its edge comes from what Inherent post-trained it to do: act as a scientific director that treats OpenAI’s Codex GPT-5.5 as a tool, much the way human researchers lean on existing software rather than writing everything themselves.
The Replica benchmark: 310 tasks, 100 papers, 36 years of science
The work is described in “Training AI Scientists to Replicate Research” (arXiv:2608.13331, submitted August 13, 2026), authored by an eleven-person team including Inherent co-founder and chief scientist Edward Hughes. To train and evaluate Faraday, the team built Replica, an automatically generated task space of 310 figure-replication tasks drawn from 100 machine learning and AI-for-science papers spanning 1990 to 2026.
Each task gives the agent a paper with the target figure redacted, a working environment with pre-installed research libraries, internet access, and limited time and compute budgets. The agent must reproduce the plot — without ever seeing it. Where an experiment cannot be completed within budget, the agent is asked to produce the most faithful scaled-down version. The split: 242 ML tasks for training, and 68 held-out AI-for-science tasks (from papers published 2012–2026) for testing.
This design deliberately targets what makes replication hard. A paper is, by definition, a lossy compression of the research that produced it — details are underspecified, and filling the gaps requires hypothesis-driven exploration rather than pattern-matching. Existing agents, heavily optimized for well-specified, closed-ended problems, tend to struggle exactly here.
Rubric-based RL: teaching “research taste” with a judge
The second technical contribution addresses the reward problem: replication quality has no unit test. Inherent’s answer is an auto-generated, per-task rubric judge — a coding agent that scores each replication attempt against criteria generated from the paper itself, aggregated over multiple samples to reduce noise. The team validated the judge against expert human rankings of rollouts (76 rankings from 19 participants), finding it agrees with human assessment and exhibits low noise — the prerequisite for using it as a reinforcement learning reward signal.
With reward in hand, Faraday was post-trained with GRPO across the 242 training tasks, with careful recipes for long-horizon stability. The result, in the paper’s own numbers:
- Faraday outperforms Claude Opus 4.8 on 73% of in-distribution ML tasks and on 60% of held-out AI-for-science tasks.
- On average, a 6% improvement over Claude Opus 4.8 and an 8% improvement over Codex GPT-5.5 on the test split.
- Optimizing Codex’s prompt to close the gap “only marginally diminishes” it — the gain comes from post-training, not prompting.
- Human experts independently rate Faraday as the stronger scientist on the rollouts where the rubric judge says it has an edge.
Crucially, the fair-comparison setup runs Faraday and an untrained Qwen3.6-27B in the same simple harness, both with Codex GPT-5.5 available as a coding tool. The two differ only in the Replica post-training — isolating exactly what the RL bought.
A director, not a coder
The most consequential finding may be qualitative. Faraday doesn’t try to out-code the frontier agents beneath it; it learns to direct them — deciding which experiments are worth running, how to scope a scaled-down version of an experiment, and when a result is scientifically meaningful. The paper’s rollout analysis concludes that Faraday “adopts a more scientifically-principled approach” than the baselines.
This is the “coding agent as a tool” (CAT) pattern: a smaller model trained to sit above frontier coding agents and supply the layer of scientific judgment they lack. Inherent explicitly frames it as a bet against harness complexity — the paper positions its results as “a stepping stone towards AI agents capable of long-horizon scientific innovation without requiring complex harnesses.”
The company tested generalization in two further ways. In a full-scale replication study — eight papers pre-filtered so each figure should be reproducible within eight hours on eight B300 GPUs — Faraday beat Claude Opus 4.8 on five of eight tasks. And on “imagined” replications (variants of real papers generated by Claude with the same claim but different datasets), Faraday’s advantage persisted, hinting that the training transfers beyond memorized task structure.
Why it matters
Three implications stand out.
First, the frontier may be an orchestration layer, not a model. A 27B model with the right post-training beat systems many times its size at a cognitively demanding task by delegating execution. If that pattern generalizes — and Inherent’s ablations suggest it isn’t reachable by prompt engineering — the competitive moat shifts from raw model scale to training recipes and task spaces.
Second, replication is a smart wedge into AI-driven science. As Inherent notes, many PhD students start their careers by replicating papers; it is the standard apprenticeship of science. An agent that can do this autonomously — at a cost far below a doctoral student — is both immediately useful (replication is a chronic bottleneck and a reproducibility-crisis remedy) and a natural curriculum step toward original discovery.
Third, the London ecosystem keeps compounding. Inherent’s dozen employees work in person in King’s Cross — the neighborhood DeepMind’s presence turned into a top AI hub. The company raised its $50M seed from Index Ventures and Radical Ventures, and Hughes is emphatic: “We believe that London is the place to be.”
Limits and caveats
The paper is refreshingly candid. The human-study sample is small (19 participants), and while participants sided with the rubric judge on 63% of disputed ranking pairs, that was higher than chance but not statistically significant (p = 0.109). The judge is itself an LLM, so some circularity risk remains. And Faraday’s scope is deliberately constrained: in silico tasks only, bounded time and compute, no physical lab equipment — though with internet access, a choice the team flags in its safety discussion.
The startup angle also warrants restraint: one strong benchmark result from a lab that just emerged from stealth is evidence of technical talent, not yet of a business. But as a signal of where AI-for-science is heading — small, taste-trained directors orchestrating giant executors — Faraday is one of the cleanest data points yet.
Inherent’s north star, per Hughes, is “building an AI scientist agent and imbuing our agents with taste.” The kind of teammate who comes back and says: “I got curious about this, and I went off and I did these experiments. What do you think of these results?”
As of this week, that teammate exists — and it fits in 27 billion parameters.
Sources
- [1] https://techcrunch.com/2026/08/22/inherent-founded-by-deepmind-alumni-says-its-ai-teammate-just-outperformed-anthropic-and-openai-at-replicating-research/
- [2] https://arxiv.org/abs/2608.13331
- [3] https://thenextweb.com/news/inherent-ai-50-million-seed-deepmind-faraday-science
- [4] https://www.indexventures.com/perspectives/inherent-designing-for-discovery/