← Back to all posts
News

Anthropic Let Claude Loose on 15 Proteins. It Failed Only One.

August 19, 2026 · 05:06 UTC · News
Anthropic Let Claude Loose on 15 Proteins. It Failed Only One.

TL;DR

Anthropic published wet-lab results on August 18 from a campaign in which Claude designed protein minibinders against 15 targets, with almost no human in the loop. Two external labs, Adaptyv Bio and Twist Bioscience, synthesized and tested the output independently. The result: 354 confirmed binders against 14 of 15 targets, out of 1,320 total designs. Between 22.6% and 35.1% of individual designs bound, against a field norm Anthropic puts at 10% to 15%. On RBX1, a target that a public design competition recently chewed on, Claude hit 40% where the entrant field managed 3.7%. This is not a benchmark score. Somebody made the molecules and they stuck.


What actually ran

Claude was not asked to be a protein design model. It was asked to operate them. Anthropic describes it orchestrating several publicly available structure design, sequence design, and co-folding models, running optimization cycles, screening candidates, and managing the workflow end to end. Think of it less as the chemist and more as the principal investigator: it never touched a pipette or invented a new folding algorithm, it just decided what to run, read the results, and decided what to run next, for two days straight, without getting bored.

Human involvement was deliberately thin. Per Anthropic, it consisted of approving certain requests Claude made (network access, code execution), fixing infrastructure problems outside the design sessions, and physically ordering the resulting designs for lab validation. Nobody was steering the science.

The campaign ran in two modes inside Claude Science, Anthropic's research workbench:

  • Multi-target: one session, all targets at once, 48 hours, up to 12,500 GPU hours total. Run with both Mythos Preview and Opus 4.8.
  • Single-target: one session per target, all running in parallel, 24 hours and up to 2,500 GPU hours each. Run with Mythos Preview.

Thirty designs were requested per target. Sixteen targets were selected; results are reported for 15 because data on one came back inconclusive.

the campaign, end to end 15 targets 1,320 designs 354 binders 14 of 15 hit 30 designs requested per target, synthesized and tested externally
One target, maltose binding protein, produced nothing that bound.

The hit rate is the story

A hit rate here means the share of individual designs that actually bound their target when made in a lab. It is the number that separates a generative model producing plausible-looking sequences from one producing molecules. Anthropic's framing is that 10% to 15% is typical for protein design campaigns today.

share of individual designs that bound in the lab Mythos, 1 target35.1% Mythos, all 1526.7% Opus 4.8, all 1522.6% field, typical10-15% source: Anthropic, validated by Adaptyv Bio and Twist Bioscience
Focus helps: one target per session beat all-15-at-once by 8.4 points.

Two things worth noticing. First, the single-target mode beat the multi-target mode by a wide margin, which is the least surprising and most useful finding in the whole post: give the agent one problem and a dedicated compute budget and it does better than when it is juggling fifteen. Anyone who has watched a coding agent fan out across a monorepo already knows this feeling.

Second, Opus 4.8 was not simply worse. It succeeded on TNF-alpha where Mythos Preview failed, and Anthropic says the obvious thing out loud: in a domain this rough, a generally less capable model can still win specific targets. Capability is not a scalar here.

The RBX1 blowout

The most legible result is RBX1, a 108-residue component of the SCF ubiquitin ligase complex and a genuinely hard, partly disordered target. It was also the subject of the GEM and Adaptyv RBX-1 Binder Design Competition, a public contest that drew more than 180 submissions from the computational protein design community and tested them in the same wet lab.

RBX1 hit rate: Claude vs the contest field Mythos, 1 target40% contest entrants3.7% GEM and Adaptyv RBX-1 competition drew 180+ submissions
Claude's top-ranked RBX1 design also outperformed the winning entry.

Across the campaign, Anthropic reports high-affinity binders against at least six targets, and binders that match or exceed the best reported affinity against at least four. Some of the strongest designs bound several times more tightly than the best previously published result. It also produced 15 fold-diverse binders containing beta-sheets, which matters because most de novo minibinder work leans heavily on helical bundles.

Where it failed

One target produced nothing: maltose binding protein. MBP is the lab mule of molecular biology, the fusion tag welded onto half the constructs in every freezer on the floor, and Claude could not get a grip on it. BBF-14, a synthetic beta-barrel, went nearly as badly: three independent binders with modest affinities and not much else.

These are not embarrassing failures so much as a useful reminder that the 22% to 35% headline is an average over a target set someone chose. Fifteen targets is a real experiment and a small one.

The other half nobody is quoting

Buried under the protein result is an analytical chemistry experiment that builders should probably care about more. Anthropic gave Opus 5 raw instrument files, NMR and LC-MS, and asked it to characterize a compound's identity and purity. It converted raw data into calibrated spectra, counted molecular features, ran quality checks, and produced reports with chromatograms and mass spectra. NMR took 23 minutes, LC-MS 19, running in parallel.

The accuracy: hydrogen counts within 0.08 units of the contract lab's own analysis, and purity of 96.4% against the lab's 96.33%. It also reverse-engineered an undocumented vendor file format on the way, which is the single most relatable thing in this entire research post. And it independently proposed the same heavy-water follow-up experiment the lab had separately decided to run.

The caveats that matter

  • This is a preprint-grade claim, not a paper. Anthropic says it intends to follow up with more extensive characterization to confirm hit rates and affinity measurements. Take the numbers as reported by the party with an interest in them, mitigated by the fact that two external labs did the synthesis and testing.
  • A binder is not a drug. Anthropic states plainly that protein minibinders are not a standard therapeutic modality, and that a high-affinity binder is the first step of a very long road.
  • Dual use is the whole subtext. Anthropic calls agentic biological discovery dual-use and notes that life science research tasks are currently blocked in its most capable model. That is the trade being made in public: the capability is real enough that the company is gating it.

Why a builder should care

Strip the biology and the shape is familiar. An agent was given a hard problem, a set of specialist tools it did not train on, a hard wall-clock deadline, a metered compute budget, and permission to iterate without supervision. The outputs were then graded by physical reality rather than by an LLM judge or a held-out test set, which is a grading rubric that does not care how confident the model sounded.

That is the interesting part. The single-target result also quietly reprices how you should think about agent scaffolding: same model, same tools, same 15 problems, and simply not making it context-switch bought 8.4 points of real-world success rate.

Key Takeaways

  • Claude designed 1,320 protein minibinders and 354 of them bound in the lab, hitting 14 of 15 targets, with synthesis and testing done externally by Adaptyv Bio and Twist Bioscience.
  • Design-level hit rates ran 22.6% (Opus 4.8, multi-target) to 35.1% (Mythos Preview, single-target), against a stated field norm of 10% to 15%.
  • On RBX1, Claude hit 40% where a 180-plus-entry public competition field hit 3.7%, and its top design beat the contest winner.
  • Dedicating one session per target beat running all 15 in one session by 8.4 points, a scaffolding lesson that generalizes well past biology.
  • It failed outright on maltose binding protein and barely scraped BBF-14, so the averages hide real per-target variance.
  • Separately, Opus 5 processed raw NMR and LC-MS instrument data to 96.4% purity against a lab's 96.33%, in under 25 minutes.

Sources: Anthropic Research: How Claude is accelerating protein design and analytical chemistry, Anthropic: Claude Science, Proteinbase: GEM x Adaptyv RBX1 competition, GEM Workshop: RBX-1 Binder Design Competition 2026, Adaptyv Bio, Twist Bioscience

AIAnthropicClaudeProtein DesignBiologyAgentsDrug DiscoveryResearch
CONSOLE
$