An AI agent ran a protein design pipeline unaided and hit 14 of 15 wet-lab targets
Anthropic reports hit rates of 22.6% to 35.1% against an industry norm of 10-15%, and says the work stays restricted in its public models pending a trusted-access programme.
Original cover art, generated for this story. THE VISSION does not republish third-party press imagery.
- Anthropic reported that Claude models designed protein binders that bound successfully against 14 of 15 targets tested in a wet lab, producing 354 confirmed binders from 1,320 designs.
- Running against all targets at once in a 48-hour session, Mythos Preview hit 26.7% and Opus 4.8 hit 22.6%; given 24 hours per individual target, Mythos Preview reached 35.1%, against the 10-15% Anthropic calls typical.
- Against one target, RBX1, the model reached a 40% hit rate versus 3.7% among entrants in an external competition, and its best design outperformed the winner of 245 submissions.
- Anthropic cautions that minibinders are not a standard therapeutic format and that a high-affinity binder is only the first step toward a drug-like molecule.
Anthropic published results on August 18 from a protein design campaign in which its Claude models drove the computational pipeline without human scientific guidance during execution. Across 15 targets, the models generated 1,320 designs — 30 per target per arm — of which 354 were confirmed as binders in wet-lab testing. Fourteen of the 15 targets yielded at least one working binder.
The headline numbers depend on how the work was scheduled. Given 48 hours of wall time and up to 12,500 NVIDIA H100 hours to attack every target simultaneously, Anthropic's Mythos Preview scored an overall hit rate of 26.7% and Opus 4.8 scored 22.6%. Narrowed to one target at a time, with 24 hours and up to 2,500 H100 hours each, Mythos Preview rose to 35.1%. Anthropic puts the typical rate for protein design campaigns at 10% to 15%.
The most striking single result came against RBX1, a protein involved in the targeted destruction of regulatory proteins. In single-target mode the model reached a 40% hit rate where competition participants had managed 3.7%, and its top design bound more tightly than the winning entry among 245 submissions. The physical work — synthesising and testing the sequences the models proposed — was carried out by Adaptyv Bio and Twist Bioscience rather than by Anthropic.
Anthropic is unusually direct about the limits. "Protein minibinders are not a standard therapeutic modality for drugs," the company writes, adding that even for established formats such as monoclonal antibodies and small molecules, "designing a high-affinity binder is just the first step in the process of generating a drug-like molecule." The company also says protein design and related biological capabilities remain restricted in the publicly available Claude Fable 5 while it builds trusted-access mechanisms and matching safeguards.
The interesting claim here is not the hit rate but the autonomy: a general-purpose model orchestrated a specialist computational pipeline well enough to beat the people who entered a competition designed for exactly that task. For a biotech team, that reframes the build-or-buy question — the bottleneck moves from having in-house design expertise to having wet-lab throughput to test what the agent proposes. The restriction on public models is the other half of the story: the capability exists and is gated, so access, not capability, is what a lab will be negotiating for.