A 27B 'AI Scientist' That Directs GPT-5.5: Inside Inherent's Faraday and the Replica Benchmark
London lab Inherent shows a 27-billion-parameter agent post-trained with long-horizon RL can out-replicate Claude Opus 4.8 and GPT-5.5 — by learning to direct the very frontier models it beats, on a 310-task benchmark where the reward is scientific taste, not code.
In 1821, Richard Phillips asked Michael Faraday — a self-taught experimentalist with little formal education — to write a review of the emerging field of electromagnetism. Faraday chose to replicate past results by hand, and in the candlelit basement of the Royal Institution he discovered that a current-carrying wire would rotate around a magnet. “Very satisfactory,” he wrote in his journal. He had just invented the electric motor.
Two centuries later, a London AI lab named Inherent has borrowed both the name and the method. On August 14 the company published research introducing Faraday, a 27-billion-parameter “AI Scientist” agent that, on the task of replicating published research, outperforms Anthropic’s Claude Opus 4.8 and OpenAI’s GPT-5.5 — models several orders of magnitude larger. A week later, mainstream coverage caught up when TechCrunch profiled the company behind it, and the results deserve a closer look than the headlines suggest.
What Faraday actually does
The task is precise: given a research paper, reproduce one of its figures — within a limited time and compute budget, and without access to the original plot. No answer key. The agent must read the methods section, design an experiment, write the code, run it, and produce a faithful reconstruction of the published result.
To make this trainable at scale, Inherent introduced Replica, a benchmark of 310 such tasks drawn from 100 machine-learning and AI-for-science papers, spanning natural language processing, materials science, structural biology, weather forecasting, and meta-learning. It is a deliberately underspecified, long-horizon challenge — precisely the shape of problem where frontier agents have historically struggled.
The numbers are striking. Faraday outperformed Claude Opus 4.8 on 73% of Replica tasks in-distribution and maintained its lead on held-out AI-for-science papers — research the base model had never seen during pretraining. The advantage was most pronounced in meta-learning, structural biology, and materials science. Notably, Faraday struggles less than frontier baselines with recent research, a hint that it learned transferable scientific skill rather than memorizing the literature.
The counterintuitive architecture
Here is the part that should reframe how you think about agent design: Faraday uses GPT-5.5 Codex as a tool.
The 27B “scientist” doesn’t write much code itself. Like a human postdoc who leans on a coding agent, Faraday decomposes the problem, decides what experiments are worth running, and delegates implementation to OpenAI’s frontier coder. Remarkably, Inherent showed that Faraday can generalize to directing a more capable agent at test time than the one it was trained with — it adapted to GPT-5.5 Codex after training only against GPT-5.4-mini.
In other words, the small model isn’t competing with the frontier model; it’s managing it. And the combination beats the frontier model working alone. As coding agents continue to improve, Inherent argues, the value of the layer sitting above them — scientific judgment — only increases. It’s an early, concrete data point for the “orchestration beats scale” thesis that has been quietly building across the agent ecosystem this summer.
Teaching taste with rubrics, not rewards
The deepest technical contribution is how Inherent handled a problem central to all of AI science: you cannot verify open-ended research automatically. A pixel-perfect plot reproduction can still be bad science — poor experimental design, unfaithful interpretation, wasted compute. Good replication requires what the researchers call “research taste.”
Inherent’s solution was an LLM judge scoring against auto-generated per-task rubrics, validated against a human study of expert preferences. Raw LLM judging is notoriously noisy — the stochasticity of the judge model makes the reward signal jittery, which destabilizes long-horizon RL. Per-task rubrics reduced that noise. Two further train-time modifications — multi-sample aggregation and turn-level credit assignment — tamed the instability that typically derails weeks-long training runs.
The result is an agent trained via long-horizon RL to internalize scientific rigor, without a hand-coded evolutionary harness and without any test-time reward. Earlier “AI Scientist” systems (the Sakana-style lineage) depended on externally programmed search loops; Faraday’s discovery behavior is, in Inherent’s framing, intrinsic — learned rather than scripted.
From replication to innovation
Replication may look like a stepping stone, and Inherent is explicit that it is one. Papers describe what worked, never the negative results that got the authors there. To replicate, an agent must recover the “99% perspiration” that never appears on the page — which requires the same hypothesis-driven exploration that produces genuinely new science.
The Replica task space is designed to escalate: remove more features from the papers given to agents, tighten or relax resource constraints, and eventually hand agents imagined papers — leading the same model, the company says, “to innovate without knowing it.”
The company behind it
Inherent emerged from stealth in May 2026 with a $50 million seed round led by Index Ventures, with Radical Ventures participating. The King’s Cross, London lab was founded by four ex-DeepMind researchers — Edward Hughes (chief scientist), Louis Kirsch, Kaloyan Aleksiev, and Tantum Collins, who previously worked on AI policy in the Biden White House. It operates as a public benefit corporation, is roughly a dozen people today, and plans to reach 20–25 by year-end — an appealing landing spot amid the post-reorganization churn at DeepMind itself.
Hughes frames the goal as an AI teammate of a specific kind: not one that tells you what you want to hear, but one that comes back and says, “I got curious about this, and I went off and I did these experiments. What do you think of these results?”
The caveats that matter
Replica is Inherent’s own benchmark, run and scored by Inherent. No independent replication of the Claude Opus 4.8 or GPT-5.5 comparisons has been published, and the paper is 47 pages of methodology that assumes the authors’ judge captures “expert taste.” The obvious question — how much of Replica’s 100 source papers could already sit inside the Qwen 3.6 base model’s pretraining data — is partially addressed by the held-out paper results, but a third-party rerun on a public benchmark like MLE-Bench would settle far more.
Still, the direction is right, and the architecture is the story: a small open-weight model, post-trained for judgment, orchestrating frontier coders it demonstrably outperforms when they work alone. If the pattern holds, the scarce resource in AI-assisted science won’t be coding ability. It will be knowing what’s worth trying — and the 2026 batch of AI scientists is starting to learn exactly that.
Sources
- [1] https://inherentlabs.ai/research/training-to-replicate
- [2] https://arxiv.org/html/2608.13331
- [3] https://techcrunch.com/2026/08/22/inherent-founded-by-deepmind-alumni-says-its-ai-teammate-just-outperformed-anthropic-and-openai-at-replicating-research/
- [4] https://aiweekly.co/alerts/inherent-says-faraday-tops-claude-gpt-55-at-paper-replication