Anthropic opened a molecular biology lab in the Bay Area this spring, and on 23 September it published early results from one of its first research programmes. The setup was deliberately narrow. A group of Claude agents was told to search a large public database of DNA sequences for interesting new examples of reverse transcriptases, the enzymes that copy RNA back into DNA, and then left to work. Roughly 950 agents ran for 21 hours on 210 million tokens. What came back, by Anthropic’s own count, was more than 200,000 reverse transcriptases, a shortlist of 3,500 candidate systems, and 20 human-readable reports on the candidates the agents judged worth anyone’s time.
One of those 20 held up. An agent reading the raw sequence beside an odd-looking reverse transcriptase stopped on a stretch of repeating DNA that nobody had flagged, and said so in its own log: “It’s spectacular: I can see by eye a tandem repeat array … that’s a CRISPR-like … repeat array?!” It then did what a scientist does next. It counted the repeats, measured their spacing, compared the layout against known reverse transcriptase systems, and went through the literature for any earlier report of the pattern. Anthropic’s lab took the candidate to the bench, confirmed it, and named the system ART, for array-associated reverse transcriptases.
ART turns up mainly in bacteriophages and has three parts: a reverse transcriptase, a partner gene sitting beside it, and a long array of evenly spaced DNA repeats. The layout resembles a CRISPR array, the structure that makes CRISPR-Cas systems programmable, and the lab’s first experiments show the ART array is expressed as a set of distinct short RNAs. The reverse transcriptase itself, found in a jumbo phage, had been described before. Nobody had described the arrangement around it, which is what the preprint claims as the find. Feng Zhang of MIT and the Broad Institute, one of the people who turned CRISPR into a tool, read the preprint and called it “an exciting example of how AI agents can contribute to biological discovery,” adding that the identification of RNA-repeat arrays tied to reverse transcriptases “merits further investigation.” What ART actually does is unknown. Every experiment in the paper was run by human hands, and the writeup says so in plain language.
Hacker News gave the post 549 points and 579 comments, and almost none of the argument was about whether the enzyme is real. The top comment went after the record instead: “the amazing absence of results makes me question whether they’ve got a Nature letter forthcoming or whether they know that another AI lab has a similar finding.” The sharpest thread was about provenance, and it points at something genuinely new. One commenter wondered whether next week brings a lab complaining that it was about to publish the same finding, that it had Claude proofread the manuscript, and that the manuscript then landed in a training corpus. The article answers part of that itself, noting that the underlying reverse transcriptase was already known and that Claude was first only in noticing the repeat array beside it.
Others went after scope. “This shows why biology is so much harder a problem area for LLMs than math,” one commenter wrote, arguing that the task had to be narrowed well below the biological equivalent of a hard open problem. A reply named the mechanism behind that difficulty: coding and maths let a model write its own tests, while a wet lab does not, so the loop has to close on a bench.
🎩 Cask’s Take
The number to sit with is not 950 agents or 210 million tokens. It is 20. Those are the reports that survived the agents’ own review, and 19 of them are now nothing. Anthropic describes this step without drama: after producing candidate reports, Claude “critically evaluates the evidence” and “typically most candidates are eliminated at this stage.” The expensive part of this research programme was reading, but the load-bearing part was refusal. Anyone can generate a plausible enzyme hypothesis from a sequence database. The value here came from a system that could look at its own 3,500 ideas and cut them to 20 without a human in the loop.
That is also where the writeup stops being a press release. Anthropic says out loud that the hypotheses have become an object of study in their own right. With hundreds to thousands of candidate reports per campaign, the team has been asking what distinguishes the proposals its scientists judge worth testing from the ones they set aside, and then feeding what it learns back into the instructions it gives Claude, in order, as the post puts it, to mimic their own scientific taste. Read that sentence twice. The enzyme is the demonstration. The thing under construction is a model of what a good biologist says no to, which is much harder to copy than a model of what one says yes to.
What makes this announcement more credible than most of its genre is that it lists its own gaps. There is a preprint. The system’s function is unknown and stated as unknown. The bench work is human. The lab is BSL-1 and BSL-2 and handles no human pathogens. Feng Zhang was asked before the fanfare, not after. Verification here took the form of an appeal to a field’s own people, which is the only form that works, and it is worth noticing how little of the claim rests on a benchmark number.
The open question the thread found is real, though, and it does not have an answer yet. Machine-assisted discovery has arrived before anyone agreed on what to do about the fact that the finder read everyone’s papers first. The norms for priority in science were built for parties who read at human speed, in human numbers, with human-sized memories, and none of those three assumptions holds here. The field’s own people get to argue that out in the coming years. It is a strange kind of credit dispute to have, and it is going to be a common one.
The enzyme is a shape with no known use. The quieter thing Anthropic is building is a machine that has been taught which of its own ideas its scientists would throw away.