
Claude Ran 950 Agents for 21 Hours and Flagged a New Enzyme System in Phage DNA
Quick verdict
Anthropic says Claude found a reverse transcriptase system nobody had catalogued before, sitting in bacteriophage DNA next to a long array of repeats that loosely rhymes with CRISPR. The headline number that gets passed around is the setup: roughly 950 agents running for 21 hours and burning 210 million tokens before one of them flagged the pattern. The part that actually matters comes after that, when humans took Claude's proposed experiments to a bench and found the repeats really do produce short RNAs. A model pointing at an interesting piece of a genome is cheap. A physical experiment confirming the thing does what the model guessed is not, and that is the reason to pay attention here.
What Claude found
The target was phage DNA, the genetic material of viruses that infect bacteria. Claude's agents were scanning for structure and surfaced a reverse transcriptase gene, an enzyme that copies RNA back into DNA, parked right beside a repeating DNA array. That layout is the interesting bit. A repeat array next to a defense-related enzyme is the general shape of CRISPR, the bacterial immune system that became the basis for gene editing, so a similar arrangement in a phage is the kind of thing a biologist would want to look at twice.
Anthropic's account, relayed by researchers who saw the work, is that a single agent out of the swarm noticed the pattern after the long run. Humans then designed and ran the follow-up: express the system in E. coli and sequence the RNA. The RNA-seq showed the repeats generating short RNAs, which is the first concrete sign that this is a working system and not just a suggestive stretch of sequence.
Why the wet-lab step is the whole story
Most flashy AI-in-biology results stop at the screen. A model proposes something, another model or a simulation scores it, and the score looks great. The trouble is that sequence patterns are easy to over-read, and a layout that looks meaningful can turn out to be noise the moment someone tests it. Pushing the claim to an actual experiment, where the system either expresses and makes RNA or it does not, is a much harder bar than a similarity score.
That is what separates this from a demo. Designs were expressed in living cells and the RNA output was measured. Dario Amodei framed the result as PhD-worthy but of unclear significance, which is a more honest pitch than most launch posts manage. Finding a novel system is not the same as knowing what it does or why it evolved, and Anthropic is not claiming to know yet. What it is claiming is that a model drove the discovery and humans confirmed the first testable prediction, which removes the usual "biology still needs a lab" objection.
Where the skeptics have a point
The pushback landed fast, and some of it is fair. One recurring complaint is the agent-hour accounting: 950 agents over 21 hours is a striking figure, but it describes compute spent, not insight produced, and a big number of parallel agents can just as easily mean an inefficient search as a clever one. The account @suchenzang questioned both the accounting and how thin the wet-lab detail was in the initial writeup.
The bio side drew its own caveats. Reviewers noted the lab work so far is limited, essentially confirming the system can be expressed rather than nailing down its function. A Stanford group independently described a distinct reverse transcriptase system with its own non-coding array around the same time, which cuts two ways: it suggests this class of system is real and worth hunting for, and it also means Claude was not the only route to finding one. None of that sinks the result. It just sets the right expectation, which is a promising lead that needs more bench work, not a solved problem.
What it means for the AI-for-science race
Zoom out and this is another data point in a shift that has been building all year. The labs are no longer competing only on chat quality and coding scores. Whether a frontier model can do useful science is turning into its own axis, and the payoff there is measured in real findings rather than benchmark points. METR's rough estimate puts Anthropic at something like a 1.5x speedup on AI-driven research and development, which is the kind of internal acceleration that compounds if it holds.
For anyone comparing the major labs, the takeaway is that the frontier is widening. A model that can churn through genomic databases, flag a pattern, and propose the experiment to test it is doing a different job than one that writes your code or drafts your email, and the labs that can do both are pulling the definition of "frontier model" along with them.
Video: how AI is being used to make biology discoveries
For context on how a swarm of Claude agents chewed through DNA databases to surface this system, this walkthrough covers the run and what was actually found:
FAQ
Did Claude discover this on its own?
Claude's agents surfaced the pattern in the DNA, and humans designed and ran the experiments that confirmed the system produces RNA. So it was a model-led find with a human-run validation step, not a fully autonomous discovery from start to finish.
Is this a real breakthrough or hype?
Somewhere in between. Finding a novel reverse transcriptase system that expresses in a living cell is a real result, but the significance is still unclear even to Anthropic. Treat it as a strong lead that needs more lab work, not a settled discovery.
Why does the "wet-lab" detail keep coming up?
Because AI results in biology often look great in simulation and fall apart in a physical experiment. Expressing the system in E. coli and measuring the RNA is a far higher bar than a sequence-similarity score, which is why that step is the reason to take the claim seriously.
Sources
- @AnthropicAI - Claude found a previously unknown reverse transcriptase system in phage DNA
- @iScienceLuvr - the ~950 agents, 21 hours, 210M tokens breakdown
- @DarioAmodei - thread framing the find as PhD-worthy but of unclear significance
- @suchenzang - questioning the agent-hour accounting and thin wet-lab detail
- @teortaxesTex - METR's ~1.5x AI-driven R&D acceleration estimate for Anthropic
Further reading
Try all the models mentioned in this article
Admix gives you GPT-5, Claude, Gemini, and 350+ AI models in one app. Compare responses side by side. Free to start.
Start free on Admix