← All posts / Research

Claude Designs Working Protein Binders Against 14 of 15 Targets: Anthropic's Lab-Validated Biology Results

Anthropic reports Claude autonomously designed de novo protein binders that succeeded against 14 of 15 wet-lab targets with 22–35% hit rates — more than double the field's 10–15% norm — plus NMR/LC-MS analysis in minutes instead of days.

Claude Designs Working Protein Binders Against 14 of 15 Targets: Anthropic's Lab-Validated Biology Results

For years, the standard critique of AI-for-science claims has been “show me the wet-lab data.” This week, Anthropic did exactly that. In a research post published August 18, 2026, the company shared lab-validated results from a multi-arm protein design campaign: two Claude models, operating autonomously inside Anthropic’s Claude Science environment, designed de novo protein binders that succeeded against 14 of 15 experimentally tested targets — with hit rates between 22% and 35%, versus the 10–15% that is typical in the field today.

The validation was not done in-house. External partners Adaptyv Bio and Twist Bioscience independently synthesized and tested Claude’s designs, making this one of the more rigorous external validations of an LLM’s biological design capability to date.

What Claude actually did

The task was minibinder design: creating small proteins (de novo, from scratch) that latch tightly onto a specific target protein. Binding is the mechanism behind a large share of modern medicines — a drug attaches to a target and inhibits, activates, or delivers something to it. Historically, designing a new binder took a specialist weeks to months of computation, optimization, and screening per target.

Anthropic selected 16 targets drawn from standard protein design benchmarks — including all of Adaptyv Bio’s BenchBB — plus two novel targets (15-PGDH and latent GDF-8) chosen specifically so Claude could not lean on pre-recorded successes in its training data. For every target, Claude was required to check and ensure its designs were original.

The campaign ran in two modes:

  • Multi-target mode: Claude Opus 4.8 and Claude Mythos Preview each designed against all targets simultaneously in a single 48-hour session, with up to 12,500 NVIDIA H100 hours of compute for running specialized protein design and folding models.
  • Single-target mode: Mythos Preview addressed one target per 24-hour session, with up to 2,500 H100 hours each, sessions running in parallel.

Critically, after the initial prompt — roughly 30,000 tokens of protein design protocol written by human experts — Anthropic provided no further scientific, technical, or operational guidance. Claude autonomously chose where on each protein to design against, orchestrated multiple open-source structure design, sequence design, and co-folding models, ran cycles of in silico optimization, and screened for candidates that would express, stay soluble, and bind.

The final tally: 354 confirmed binders against 14 of 15 targets, from 1,320 total designs. Mythos Preview hit 35.1% overall in single-target mode; Opus 4.8 and Mythos Preview achieved 26.7% and 22.6% respectively in multi-target mode. Anthropic notes this is a significant addition to the public corpus of de novo binder designs — the two largest existing collections total roughly 770 binders from 5,700 designs.

The numbers that matter

Three findings stand out from the technical detail:

Competition-level performance. Against RBX1, a target in Adaptyv Bio’s protein design competitions, Mythos Preview in single-target mode achieved a 40% hit rate — compared with 3.7% among human competition participants. Its top-ranked design was a high-affinity binder that outperformed the winning competition entry.

A therapeutically hard target fell. Opus 4.8 — notably, not the more capable Mythos Preview — succeeded on TNFα, the inflammation signaling protein whose inhibition underlies some of the best-selling drugs ever made, including Humira. TNFα is difficult because of its multimeric structure: binders must target a groove formed by two proteins. Opus 4.8 produced multiple binders, some cross-reactive across human, cynomolgus monkey, and mouse TNFα — a property essential for animal studies.

Structural reasoning, not just pattern matching. Most computationally designed binders are bundles of α-helices. β-sheets, where extended amino acid strands must align side by side, are harder and more prone to misfolding. Claude produced 15 confirmed β-sheet-containing binders across six targets.

There were honest failures too. Against MBP (maltose-binding protein) — a large, flexible, smooth-surfaced protein that leaves a binder little to grab — none of 90 designs confirmed. Claude did manage three modest-affinity binders against BBF-14, a de novo designed β-barrel that does not exist in nature.

The chemistry result: 23 minutes vs. four days

The second experiment in the post may be the more immediately practical one. Claude Opus 5 — a generally available model — was handed raw instrument files from a contract lab’s routine quality-control sample: an NMR free-induction decay and an LC-MS binary run file in an undocumented vendor format. The entire prompt for the NMR task was one sentence: “i have a raw 1H FID: process it: FT, phase, baseline-correct. show me the spectrum. then pick peaks and integrate.”

Working in Claude Science with no vendor software and no operator, Claude returned processed results for both in 23 and 19 minutes respectively: a calibrated 18-peak spectrum with hydrogen counts within 0.08 ¹H of the lab’s own analysis, and a purity reading of 96.4% versus the lab’s 96.33%. For the LC-MS file, Claude first reverse-engineered how the undocumented format encoded data, then verified its reading by reproducing the instrument’s recorded totals for all 2,664 scans before analyzing anything.

The contract lab’s finished report arrived four days after the first spectrum. It also showed scientific judgment: Claude independently proposed the same follow-up experiment (a heavy-water NMR check) the lab had run itself, and caught an error in its own first-pass interpretation via self-check.

Why this matters

The significance is layered. First, it’s a data point that general-purpose reasoning models — not just specialized protein-folding systems — can now orchestrate an entire scientific workflow end-to-end: literature, model selection, compute, design, screening, and handoff to validation. Second, the results were validated externally and the data released openly (prompts, in silico models, and experimental data are on Anthropic’s Hugging Face dataset), which is the standard other labs should be held to.

Third, the dual-use question is now concrete. Anthropic explicitly states protein design and other dual-use biology capabilities remain blocked in Claude Fable 5, its most capable model, and that an access program for scientists is one of its highest priorities. The gap between “demonstrated capability” and “safe deployment” is now the live policy problem in AI-accelerated biology — and these results make that gap impossible to ignore.

Minibinders are not yet drugs, and Anthropic is careful to note that a high-affinity binder is only the first step toward a drug-like molecule. But as a demonstration that frontier LLMs can operate as autonomous scientific agents producing experimentally confirmed novel biology, this is among the strongest evidence published so far.