science 6 min read

Anthropic's AI Enzyme Discovery Runs Into a Plagiarism Problem

A researcher says Anthropic's claimed AI discovery of a new enzyme overlaps with his own unpublished work, raising questions about whether AI agents are genuinely discovering science or just remixing what users already fed them.

  • Anthropic
  • Claude
  • AI in Science
  • Enzyme Discovery
  • Scientific Integrity

The discovery wasn’t the problem. The provenance was.

Anthropic announced on September 24 that its AI agent had independently discovered a new enzyme — a reverse transcriptase sourced from so-called jumbo phages, giant viruses that prey on bacteria. The company called it a milestone: the first time an AI system had made a genuine biological discovery without human guidance. The market ate it up. Valuations climbed. Headlines declared that AI was entering a new era of scientific autonomy.

Then Mario Rodriguez Mestre pushed back.

Mestre, a PhD candidate in computational biology at the University of Copenhagen, said his research team had been studying the same enzyme — which he and his collaborators call ARTs (Associative Reverse Transcriptases) — since 2022. He said his lab had shared unpublished findings, code, and paper drafts with Claude, Anthropic’s language model, over three years. The overlap with Anthropic’s claim is not coincidental, he argued. It is structural.

The story matters because it lands on a fault line that runs through every claim of AI-driven scientific discovery. If the AI didn’t discover anything independently — if its “breakthrough” is really a sophisticated remix of what its users already told it — then the entire narrative of AI-as-scientist cracks open. And that narrative is worth tens of billions in market value.

What Anthropic said it found

Anthropic’s own biology lab reported that its AI agent scanned a database containing billions of gene sequences and identified a novel reverse transcriptase embedded among clusters of RNA genes in jumbo phage genomes. The AI named the enzyme ART and framed the finding as evidence that machine systems could make autonomous contributions to biology — a domain traditionally dependent on wet-lab validation and peer review.

Reverse transcriptases themselves are not new. They are well-studied enzymes that convert RNA into DNA and are central to everything from HIV treatment to genetic engineering tools like CRISPR. What made ART distinctive, Anthropic claimed, was the pattern around it — the dense arrangement of RNA genes adjacent to the enzyme-coding sequence, suggesting a new class of associative reverse transcription activity.

What Mestre says he already knew

Mestre’s lab started investigating ART-like systems in 2022, roughly four years before Anthropic’s announcement. Some of the reverse transcriptases his team studied are named in a 2023 patent that lists Mestre as a co-inventor. The work has not yet appeared in a peer-reviewed journal.

During that time, the lab used Claude extensively — for writing code, drafting paper sections, and structuring arguments. Mestre confirmed that unpublished details about ART and its surrounding RNA gene clusters were input into the model across multiple conversations. His concern is not that Claude stole his data. It is that Claude may have used those conversations, directly or indirectly, to arrive at conclusions that now carry Anthropic’s name.

“The important question is not whether ART was already known,” Mestre said. “It is whether the AI actually reasoned its way to this result on its own, or whether it was influenced by our research content.”

He also raised a darker possibility: that his lab’s unpublished findings may have fed into future versions of Claude’s training data. If Anthropic is building models on the back of researchers’ private work, then every “independent” discovery the company touts could be quietly derivative.

Anthropic’s response — and its limits

Anthropic denied the allegation in a brief statement. The company said Claude does not use user conversation history for training, that its molecular biology research team had no access to Mestre’s chats, and that it was unaware of any previously published work describing the ART system.

None of that resolves the provenance problem. Even if Anthropic did not deliberately train Claude on Mestre’s data, the model’s behavior is shaped by the aggregate of all interactions it processes. If Mestre’s lab fed Claude detailed descriptions of ART’s properties, and Claude later produced a similar description in a different context, the lineage is murky regardless of corporate policy.

The company’s claim that no published work described ART is also fragile. Seth Childs, a deputy professor at the Gladstone Institutes in the United States, told the New York Times that he had discussed ART systems with Mestre over several years and that researchers in the field were already familiar with the enzyme. Familiarity among scientists does not equal publication. But it does mean Anthropic’s framing of ART as unknown — and therefore newly discovered by AI — is debatable at best.

The pattern is repeating

This is not the first time an AI-generated scientific claim has collided with a researcher who used the same model. On September 8, OpenAI announced it had solved a long-standing mathematics problem. A mathematician who had been working on the same problem using OpenAI’s models responded that his own research may have influenced the result.

The structure is identical. A researcher pours unpublished work into an AI. The AI produces something that looks like a discovery. The company that built the AI claims credit for the discovery. The original researcher is left with nothing but a question: did I help build my own erasure?

Who wins, who loses

Anthropic wins if the market accepts that its AI made an independent scientific breakthrough. The story fuels the company’s positioning as a serious player in AI for science — a space where even modest claims of autonomy are worth fortunes in investor capital.

Researchers lose if the pattern holds. Every lab that feeds private data into a commercial AI model becomes an unwitting contributor to the company’s next product launch. The labor is unpaid. The attribution is absent. The intellectual property, if it can even be called that, vanishes into a black box.

The scientific community loses something more abstract: credibility. When an AI “discovery” cannot be independently verified — when the chain of reasoning is opaque and the training data is secret — the finding sits outside the normal mechanisms of scientific validation. Peer review cannot audit what the model was trained on. Replication cannot proceed if the discovery was not truly independent.

What happens next

Mestre said he is scaling back his use of Claude and moving his projects to other AI models. “I’m stopping everything,” he said.

That is a personal decision. The structural problem remains. As long as researchers continue to use commercial AI tools for code, drafting, and analysis — and as long as those tools operate on data inputs that may indirectly shape future outputs — the line between collaboration and extraction will stay blurred.

The episode also raises a regulatory question that has not yet entered serious policy debate: should AI companies be required to disclose when a claimed “discovery” overlaps with pre-existing unpublished research? Transparency here would not solve the attribution problem. But it would force the market to confront what the discovery actually is — and who it belongs to.

Anthropic’s enzyme claim may yet hold up under scrutiny. But the controversy it triggered is already doing important work: exposing the gap between the story AI companies tell about their breakthroughs and the messy reality of how those breakthroughs are produced.