Claude Designed Working Protein Binders for 14 of 15 Targets — No Human in the Loop
Anthropic ran Claude autonomously through full protein binder design campaigns: 354 of 1,320 designs bound in wet-lab tests, beating open competition entries on hit rate and affinity.
For decades, designing a protein that binds a chosen target has been a specialist’s craft: weeks of expert decisions about epitopes, scaffolds, and filtering thresholds, followed by days of orchestrating finicky computational tools. Anthropic has now published lab-validated results showing that Claude can do the entire job itself. Given nothing but a target’s name and a written protocol, AI agents ran complete binder design campaigns against 16 protein targets — and the proteins they designed actually worked, confirmed by two independent contract labs.
The numbers, published August 18 in Anthropic’s paper “Autonomous de novo protein binder design with Claude,” are difficult to dismiss. Of 1,320 designs synthesized and tested, 354 bound their target — a 27% hit rate against an industry norm of 10–15%. Among the designs Claude ranked first for each target, 49% bound. Binders were found against 14 of the 15 targets with interpretable measurements, spanning cytokines (TNFα, VEGF-A), cell-surface receptors (PD-L1, TREM2, TrkA), viral proteins (Nipah G, EBV’s BHRF1), the Cas9 enzyme, and the E3 ligase subunit RBX1.
How the campaigns ran
The setup was deliberately hands-off. Anthropic’s team wrote roughly 16,000 words of “protocol prompt” — the working knowledge of a binder design campaign, covering stages, open-source tools, and selection criteria — but specified no epitope, scaffold, or sequence for any target. From there, Claude Opus 4.8 and Claude Mythos Preview ran 24- to 48-hour campaigns on their own: researching each target’s biology, choosing which surface to attack, installing and running open-source design tools like RFdiffusion3, Genie 3, and Proteina-Complexa, optimizing candidates in silico, and delivering 30 ranked designs per target.
Humans intervened only at the beginning and the end: choosing targets, providing a cloud GPU account with a fixed budget, placing synthesis orders, and interpreting binding data. Every design decision in between belonged to the model. Two contract research organizations — Adaptyv Bio in Switzerland and Twist Bioscience — then synthesized every design exactly as delivered and measured binding, with neither lab seeing the other’s results.
Beating the competitions
The most striking comparisons come from targets that have been the subject of open protein design competitions. On RBX1, an E3 ligase subunit, a recent community competition saw 9 of 245 de novo designs bind. Claude’s score: 28 of 90. Its tightest RBX1 binder had a K_D of 3.9 nM — measured on the same plate, the competition’s winning entry bound at 45 nM. That is roughly a twelve-fold improvement in affinity over the best human-team effort, measured side by side.
On TREM2, Claude’s top ten designs from each of three campaigns produced 9, 10, and 10 binders. On Nipah G, Claude hit 19 of 90 against the competition’s 69 of 666. And on TNFα — the target of five approved biologic drugs and a protein where multiple prior de novo design efforts reported zero binders — Claude found twelve, including species cross-reactive designs with apparent K_D down to 0.70 nM.
Real-world signals
Several secondary results suggest these aren’t just benchmark artifacts. Cross-species reactivity, a prerequisite for preclinical development and only a secondary objective in the prompt, appeared unprompted: 130 of 233 binders tested also bound the mouse version of their target, and 154 of 179 bound the cynomolgus ortholog. Claude concentrated 91–100% of its designs on the natural binding epitope for eight of ten targets where that’s known — it worked the surfaces that evolution uses. And 98% of tested designs had no sequence match in the Protein Data Bank above 30% identity: these are genuinely novel proteins.
The failures are reported with unusual candor. Nothing bound maltose-binding protein (0 of 90). Only one design bound 15-PGDH, and three bound the artificial β-barrel BBF-14 — and in each case the co-folding confidence scores that guided the campaign gave almost no warning, scoring failed campaigns nearly as high as successful ones. Anthropic’s own conclusion: a confident computational prediction is a useful filter but not a guarantee, and experimental screening remains irreplaceable.
A same-day chemistry result
Alongside the protein work, Anthropic published a second lab result involving analytical chemistry. Claude Opus 5, a generally available model, was handed raw NMR and LC-MS instrument files from a contract lab with only a two-sentence prompt. It returned finished analyses in 23 and 19 minutes respectively — converting the raw NMR signal into a calibrated 18-peak spectrum with hydrogen counts within 0.08 ¹H of the lab’s own reading, and measuring purity at 96.4% versus the lab’s 96.33%. What normally occupies a trained chemist for hours of instrument-software drudgery was compressed to minutes.
The safety question
The subtext matters as much as the results. Anthropic notes that life-science research tasks are currently blocked in its most capable model, and that launching an access program for scientists is “one of our highest priorities.” The published results relied on a combination of Mythos and Opus models, with the frontier capability deliberately fenced off while a vetted access pathway is built. Every tool Claude used in the campaigns is open-source, which cuts both ways: it democratizes the technique for any laboratory, and it underscores why Anthropic is moving carefully on who gets the most capable version. The full prompts, computational models of all 1,440 designs, and binding data for all 1,320 tested designs were released on Hugging Face under CC BY 4.0 and MIT licenses.
Why it matters
The significance is not that AI can design a binder — specialized pipelines have done that for years. It is that a general-purpose AI agent, given a written protocol and no per-decision help, exercised the full expertise of a protein engineering team across 16 targets simultaneously, within one to two days per campaign. Anthropic explicitly declines to claim Claude beats expert humans with the same tools. What the study establishes is autonomy and completeness: the same frozen protocol served all 16 targets, and everything — including failures — was synthesized, measured, and published. For laboratories with targets of interest but no computational protein design group, the barrier to entry just dropped from a team to a prompt.