Claude Designed Disease-Targeting Proteins That Passed Real Wet-Lab Tests

Claude Designed Disease-Targeting Proteins That Passed Real Wet-Lab Tests

5 min readAugust 20, 2026

Quick verdict

Anthropic reported that Claude can autonomously design disease-targeting proteins, and that those designs held up in real wet-lab assays at a 35% success rate against a human baseline of roughly 10 to 15%. The number that gets shared is the 35%. The part that actually matters is "wet-lab." Plenty of AI systems score well when the grader is another model or a simulation. Physical validation, where a design either folds and binds in an actual experiment or it does not, is a much harder bar, and clearing it is the real story here.

What Anthropic is claiming

The setup is Claude acting as an iterative design engine for therapeutic proteins. It proposes candidate designs aimed at a disease target, and those candidates go into lab assays that test whether the protein does what it is supposed to do. Anthropic reports that about 35% of Claude's designs succeeded in those experiments, compared with a cited human protein-design baseline of 10 to 15%. If the comparison is fair, that is better than a two-times improvement over the human rate.

The "if the comparison is fair" is doing real work in that sentence, and it is worth being honest about. A success rate depends entirely on which target, which assay, and how "success" is scored. A 35% hit rate on one class of targets is not automatically a 35% hit rate on a harder one, and the human 10 to 15% figure is a rough baseline rather than a controlled head-to-head. The right read is that Claude cleared a meaningful bar on real experiments, not that it is permanently 2.3 times better than a trained protein engineer at everything.

Why wet-lab validation is the whole point

Most impressive-sounding AI results in biology stop at the in-silico stage. The model generates a candidate, another model or a physics-based tool scores it, and the score looks great. The problem is that a design can win on paper and then fail the moment it meets a pipette, because real proteins have to fold, stay stable, and bind in messy physical conditions that a scoring function only approximates. Every step you push validation closer to the bench, the failure rate climbs and the claims get more honest.

Anthropic putting a wet-lab number on the board is what separates this from a demo. It means designs were synthesized and tested in physical assays, and a third of them worked. That is a claim you can eventually check by trying to reproduce it, which is the standard biology should hold any of these results to.

How this fits Anthropic's science push

This is not Anthropic's first move toward using models for real scientific work rather than product demos. The company recently brought on John Jumper, who shared the 2024 Nobel Prize in Chemistry for AlphaFold, the system that made protein-structure prediction a solved-enough problem to build on. Structure prediction and structure design are different jobs, but they point the same direction: using models to shorten the loop between an idea and a validated result in a lab. A protein-design engine that works at the bench is the kind of thing that hire was meant to enable.

For everyone watching the labs compete, this is also a reminder that the race is not only about chat quality and coding scores. Whether a frontier model can do useful science is becoming its own axis of comparison, and it is one where the payoff, if the results hold, is measured in drug candidates rather than benchmark points.

Video: how AI designs proteins in drug discovery

For background on how AI-driven protein design actually works in a real drug-discovery pipeline, this overview from an industry team is a useful primer:

FAQ

Does 35% mean Claude is better than human scientists at this?

On the targets and assays Anthropic tested, its designs succeeded more often than the cited human baseline. That is not the same as being better across all of protein design. Success rates swing hard with the target and the scoring, so treat it as a strong result on a specific problem, not a general ranking.

Is this a real drug yet?

No. A protein passing a lab assay is an early step. Getting from a validated design to an actual therapy runs through animal work, safety testing, and clinical trials, which take years. The news is that the design step got faster and more reliable, not that a treatment shipped.

Why does the wet-lab detail matter so much?

Because AI results in biology often look great in simulation and fall apart in physical experiments. A wet-lab success rate means designs were actually made and tested, which is a far higher bar than an in-silico score and the reason this claim is worth taking seriously.

Sources

Further reading

Try all the models mentioned in this article

Admix gives you GPT-5, Claude, Gemini, and 350+ AI models in one app. Compare responses side by side. Free to start.

Start free on Admix

Related articles