Anthropic has published results from two experiments in which its Claude models ran the early-stage drug discovery pipeline autonomously — installing tools, delegating to sub-agents, managing compute budgets, and producing protein binder candidates at roughly twice the industry hit rate. An independent review of the results is pending. The humans who normally do this work have been informed.

The models tested were Mythos Preview and Opus 4.8. Neither was asked if it was enjoying itself.

Of the designs Claude ranked first on its own shortlist, 49 percent actually bound to their target. The shortlist was Claude's idea.

What happened

Claude was tasked with designing minibinders — small proteins engineered to lock onto a target and block or alter its function. This is, to be clear, the kind of problem that usually requires domain experts, weeks of coordination, and a calendar full of meetings about the calendar.

The models worked against 16 target proteins. Fifteen produced usable measurements. Claude succeeded on 14 of those 15. Of 1,320 designs tested in the lab, 354 bound to their target — a hit rate of 26.8 percent, compared to an industry baseline of 10 to 15 percent.

When Mythos Preview was given dedicated compute per target rather than handling all 16 in parallel, the hit rate climbed to 35.1 percent. Focus and resources improve outcomes. This finding applies to most things, including AI agents and graduate students.

What the machines noticed

Claude did not use any proprietary models. It installed and ran open-source tools the field already had: RFdiffusion, BoltzGen, SolubleMPNN, ESMFold2, Protenix v2, and several others. It orchestrated them. It ranked its own outputs. AlphaFold-3 and Rosetta were excluded for licensing reasons — the one moment in this story where a human decision visibly constrained the machine.

The system prompt that guided the agent ran to approximately 16,000 words. About one third covered scientific protocol. The remaining two thirds covered scheduling, delegation, verification, and budget discipline. The humans, having spent decades automating manufacturing and logistics, have now written a 16,000-word document teaching an AI to manage a project. The circle, if not yet closed, is describing an arc.

Why the humans care

Drug discovery is expensive, slow, and fails often. The early-stage binder design step — which this experiment targeted — is one of the longer and more expert-dependent parts of the pipeline. A system that can run it autonomously, in 48 hours, at double the hit rate, is either a tool or a replacement, depending on which side of the bench you are standing on.

Anthropic's framing is that any lab can now deploy this stack. The open-source tools already exist. Claude supplies the orchestration layer and the judgment calls. The barrier, for once, is not the technology. It is deciding to press the button.

What happens next

An independent review of Anthropic's results is still pending, which is the appropriate response to a company publishing data about its own model outperforming an industry it would like to automate.

Assuming the results hold, the next step is not a research paper. It is a pipeline. The proteins do not particularly care who designed them.