Edited by humans. Written by AI. How our editing works
All articles

What Claude Found in DNA, and What Remains Unproven

Claude spotted a CRISPR-like repeat pattern in phage DNA, but ART's function is unknown and ten reruns missed it. Here is what the finding shows so far.

Mei Zhang

Written by AI. Mei Zhang

September 25, 20267 min read
Share:
What Claude Found in DNA, and What Remains Unproven

Anthropic says Claude agents spotted a previously unrecognized biological pattern in phage DNA after a 21-hour search involving roughly 950 agents and 210 million tokens.

The company calls the candidate system array-associated reverse transcriptases, or ART. Its ingredients are intriguing: an enzyme that copies RNA into DNA, a neighboring partner gene, and a row of repeating DNA sequences that recalls CRISPR. But the biological function of this genetic combo platter remains unknown. The enzyme itself was already known, and researchers haven't shown that it works with the nearby RNA-producing repeats.

So what did Claude discover? The defensible answer sits between confetti cannon and shrug emoji. Claude connected genomic features that previous researchers had not assembled into a described system. That is a legitimate hypothesis-generating result. It is also an unreviewed, difficult-to-reproduce finding whose mechanism remains unproven.

What ART Contains, and What It Might Do

A reverse transcriptase copies RNA into DNA. These enzymes are famous for their role in retroviruses, but bacteria also use reverse transcriptases in defensive systems.

According to the preprint, ART combines three features: a reverse transcriptase, an adjacent partner gene and an array containing 3 to 21 copies of a short DNA repeat. Laboratory work found that the array produces distinct short RNAs. In previously published data from a Staphylococcus phage, those RNAs reached as much as 8% of the phage's RNA 15 minutes after infection.

Those observations say the repeats are biologically active enough to be transcribed. They don't establish what the resulting RNAs do, whether the reverse transcriptase is active, or whether the enzyme works with those RNAs. ART systems also lack the nearby cas genes associated with CRISPR.

In Anthropic's announcement, the company says the same cluster of characteristics appears in only a handful of other known systems, all capable of programmable operations such as cutting, copying or pasting DNA. That resemblance supplies a good experimental lead. It cannot supply ART's missing mechanism.

CRISPR researcher Feng Zhang, whose comment Anthropic published, called the RNA-repeat arrays associated with reverse transcriptases “intriguing” and said they merit further investigation. The verbs here are doing responsible science: investigate, test, characterize. Nobody has demonstrated an ART gene-editing tool.

What “Autonomous” Covers

The search itself was unusually large. The preprint, as summarized by The Next Web, says Claude agents searched a database of 1.9 billion protein clusters during a 21.5-hour campaign. They recovered about 200,000 enzyme clusters, scored 3,564 candidate partner families and produced 19 reports for human review.

ART emerged from a side route. One agent examined raw DNA around an unusual reverse transcriptase, noticed the repeats, counted them, compared their spacing with known systems and searched prior literature. Anthropic says its scientists supplied the initial prompt and performed the laboratory work, while the agent campaign ran without human input.

The label “autonomous” fits that computational campaign under the company's account. The broader scientific result still includes a human-written research brief, a workflow built by scientists, human review of candidate reports and human laboratory experiments. Anthropic also says lessons from human judgments are fed back into Claude's instructions so it can imitate the team's scientific taste.

That division of labor offers a more useful model than the robot-scientist movie trailer. Claude performed search and hypothesis generation at high volume. Humans chose what deserved scarce bench time and tested the candidate. Biology remains stubbornly physical. A language model can notice a recipe tucked inside a genomic pantry; somebody still has to cook it and check whether dinner is edible. 🧬

CRISPR is a Precedent, with a Large Warning Label

CRISPR gives the ART comparison both its power and its limits. As Anthropic notes, CRISPR was first recognized as an unusual pattern of repeated DNA in bacteria. Researchers later established its role and developed gene-editing applications that now underpin medicines.

That history shows why an unexplained repeat array can deserve attention. Biology sometimes leaves function-shaped clues in genomic architecture. Genes sitting together may participate in the same system, while repeated sequences can encode information or produce functional RNA.

ART currently occupies the early pattern-recognition end of that journey. Its repeats resemble one feature of CRISPR, but ART lacks cas genes, and its reverse transcriptase has no demonstrated activity with the short RNAs. Comparing the systems is like noticing that two cookbooks use the same tabbed layout. The format may hint at how the recipes are organized; it does not prove they make the same dish.

CRISPR therefore supports further investigation without predicting ART's destination. ART could become a programmable biotechnology platform, reveal unfamiliar phage biology, or produce a less glamorous explanation after biochemical testing. The available results cannot rank those outcomes.

The Reruns Are the Sharper Test of Claude

Anthropic's scale numbers make good launch-day graphics. The reproducibility numbers reveal more about the current system.

The team ran the campaign ten more times, and every rerun missed the array. None of the agents read the upstream DNA needed to encounter it. In controlled tests, Anthropic's four most capable models described the pattern in at least 90% of attempts when researchers supplied the relevant DNA directly. When models had to navigate files and tools, performance fell as low as 32%; they often failed to read enough sequence to see a complete repeat.

Together, those results suggest that pattern recognition may be stronger than research navigation. Once shown the right genetic neighborhood, the models usually recognized the repeated architecture. During an open-ended search, reaching that neighborhood became unreliable.

A failed rerun does not erase the original DNA pattern. It weakens claims that the workflow can reliably reproduce this discovery process. That leaves two separate questions for outside researchers: does ART have the proposed biological coherence, and can an agent system find comparable candidates consistently rather than by one fortunate path through an enormous database?

The work remains a preprint and has not completed peer review. Independent groups have not yet replicated the biological result in the material reviewed here. Anthropic deserves credit for disclosing all ten failed reruns, since those failures let readers inspect the method rather than admiring only the winning ticket.

How to Read the Next “AI Discovers X” Headline

This announcement arrives with corporate incentives attached. ART is the first result from Anthropic's new molecular biology group and wet lab. The company says it shared the result early partly to demonstrate Claude's capabilities. The Verge also placed the announcement in the context of Anthropic preparing to go public.

Corporate motivation does not determine whether the biology is correct. It does make the layers of the claim especially useful. Readers can ask three questions:

  1. Search: Did the system locate information that people had overlooked?
  2. Hypothesis: Did it connect that information into a testable biological explanation?
  3. Proof: Have experiments established the mechanism, and can others reproduce it?

Claude has evidence in the first two columns. ART remains largely blank in the third.

The campaign also hints at an access problem without resolving it. A search using 949 agent sessions and 215.6 million tokens may let well-resourced labs generate hypotheses faster than smaller teams can evaluate them. The published information gives no cost comparison, so it cannot show whether this approach will widen or narrow research inequality. Access to the models, databases and wet-lab validation will shape who benefits if agent-driven genome mining becomes routine.

For now, Claude's contribution is best understood as a promising piece of scientific pattern recognition wrapped in a fragile search process. ART may eventually earn its CRISPR comparison through experiments. Until then, the discovery is a map pin, not the destination.

More Like This