← All posts / Research

Claude Designs Protein Binders Better Than Human Experts: Anthropic's 14-of-15 Campaign and a 23-Minute Chemistry Workflow

Anthropic reports that Claude designed protein binders against 14 of 15 targets with hit rates up to 35% — roughly triple the human-typical 10-15% — and processed raw NMR and LC-MS instrument files in minutes, matching a contract lab's own analysis.

Claude Designs Protein Binders Better Than Human Experts: Anthropic's 14-of-15 Campaign and a 23-Minute Chemistry Workflow

On August 18, 2026, Anthropic published a research post that may mark one of the clearest demonstrations yet that frontier AI models can do useful laboratory science end-to-end. In two experiments, Claude designed novel protein binders against real drug targets — beating top human experts on hit rate and binding affinity — and then, in a completely different discipline, took raw instrument files from a contract chemistry lab and returned finished analytical results in minutes rather than days.

The significance is not just the numbers. It is that both tasks were executed autonomously, inside Anthropic’s Claude Science workbench, with minimal human involvement, using the same general-purpose reasoning capabilities that power Claude’s everyday chat and coding products.

Experiment 1: A protein design campaign, run by an agent

Designing a protein binder — a small protein engineered to latch tightly onto a disease-relevant target — has historically taken specialist teams weeks or months of computation, optimization, and wet-lab screening per target. Machine-learning structure models have sped this up, but they still require days of expert orchestration.

Anthropic gave Claude (in Mythos Preview and Opus 4.8 variants) a roughly 30,000-token protein design prompt, access to the internet, a corpus of protein design literature, connectors for Google Drive, Slack, Gmail, and BioRxiv, and GPUs to run specialized folding and design models — up to 12,500 NVIDIA H100 hours in a 48-hour multi-target session. Then the humans stepped back.

Claude did everything else on its own: it chose where on each target protein to design against, generated candidate structures and sequences by orchestrating multiple open-source design and co-folding models, ran cycles of in-silico optimization, and screened candidates for novelty, solubility, and expressibility. For each of 15 targets it produced 30 candidate binders, which were then synthesized and tested by two independent external labs, Adaptyv Bio and Twist Bioscience.

The results, validated in the wet lab:

  • 14 of 15 targets successfully bound. In total, 354 confirmed binders emerged from 1,320 designs.
  • Hit rates of 22.6% to 35.1%, depending on model and mode — versus the 10-15% that is typical in protein design campaigns today. Mythos Preview hit 35.1% when focused on single targets in parallel 24-hour sessions.
  • High-affinity binders (KD below 10 nM) against at least six targets, with binders matching or exceeding the best previously published affinity against at least four.
  • Against RBX1, a target from Adaptyv Bio’s design competition, Claude achieved a 40% hit rate where human participants averaged 3.7% — and its top-ranked design outperformed the competition’s winning entry.
  • Against TNFα — the inflammation target behind Humira, one of the best-selling drugs ever, and one that multiple expert groups have struggled with — Claude Opus 4.8 designed binders that worked across species, binding human, cynomolgus monkey, and mouse TNFα. Notably, the generally less capable Opus 4.8 succeeded where Mythos Preview failed, a reminder that capability in scientific domains does not track raw model rank neatly.
  • Claude also produced 15 confirmed binders containing β-sheets, a protein structure that is notoriously harder to design than the α-helix bundles most computational tools default to.

The campaign wasn’t flawless. Against maltose-binding protein (MBP), a large, flexible target with a smooth, water-loving surface, none of 90 designs confirmed. Against BBF-14, an artificial β-barrel that exists nowhere in nature, Claude managed only three modest-affinity binders. Anthropic is candid about both the failures and the need for further characterization.

For scale: the two largest public collections of de novo designed binders — proteinbase.com and the Overath et al. corpus — together contain about 770 binders from 5,700 designs against 40 targets. Claude’s 354 binders from 1,320 designs represent a material expansion of the public design corpus in a single campaign, and Anthropic has released all prompts, in-vitro data, and in-silico data.

Experiment 2: Analytical chemistry in 23 minutes

The second experiment targeted a different bottleneck: the tedious, expert-heavy work of interpreting analytical instruments.

Every time a chemist synthesizes a molecule, they must confirm identity and purity using NMR spectroscopy and LC-MS. The instruments produce raw files in proprietary vendor formats, and a chemist then manually matches each spectral peak to atoms in the proposed structure — typically 30 to 60 minutes per sample, with lab reports often lagging days behind.

Anthropic gave Claude Opus 5 — a generally available model, not a research preview — nothing but a contract lab’s raw files and a two-sentence plain-language prompt. No vendor software, no operator.

Claude processed both files in parallel, returning finished NMR results in 23 minutes and LC-MS results in 19 minutes:

  • Its hydrogen counts per peak were within 0.08 ¹H of the lab’s own analysis.
  • Its purity measurement came in at 96.4% versus the lab’s 96.33%.
  • For the LC-MS file, written in an undocumented vendor format, Claude worked out the encoding itself, then verified its parsing by reproducing the instrument’s own recorded totals for all 2,664 scans before analyzing anything — a genuine act of scientific self-checking.
  • It flagged four ambiguous NMR peaks and proposed adding heavy water to confirm them — the exact follow-up experiment the contract lab had independently run three days earlier.
  • When given the heavy-water data, Claude caught and corrected an overstatement in its own first reading, concluding — correctly — that only two of the four flagged peaks had fully disappeared.

The lab’s finished written report for that sample arrived four days after the first spectrum was acquired. Claude delivered an equivalent report, with reusable code and its own list of caveats about measurement trustworthiness, inside 25 minutes.

The safety question

Capabilities like these are unambiguously dual-use: the same autonomous biological reasoning that designs therapeutics could, in the wrong hands, enable dangerous research. Anthropic notes that protein design and other dual-use biology capabilities remain blocked in its most capable model, Fable 5, and that life-science research tasks are gated while the company builds “trusted access programs.” A formal access program for scientists is described as one of the company’s highest priorities, with details promised soon.

Why it matters

Taken together, the two experiments sketch the outline of an AI-accelerated drug pipeline: autonomous molecular design at the front end, automated analytical verification in the middle. Anthropic is explicit that a binder is only the first step toward a drug, and that many bottlenecks in drug development are policy and operations problems, not science problems. But compressing a weeks-long design campaign into a 48-hour autonomous run — and a days-long analytical workflow into 23 minutes — changes the economics of early-stage discovery in a way few would have predicted even a year ago.

The company says it is now extending the work toward running the entire development process end-to-end, across all drug modalities. For an industry where the average drug takes a decade and over a billion dollars to reach market, that is a claim worth watching closely.