Data & Training 24 September 2026 7 min read 1,568 words

Taste is now a training set

Anthropic says Claude found a new enzyme system in phage DNA after 21 hours, roughly 950 agents and 210 million tokens. The number that matters more is the one just after it: 3,500 candidate systems became 20, and the rule that did the cutting is what the company says it is now studying.

The argument

The scarce resource in genome mining has moved from candidates to the judgment that throws candidates away, and the loop Anthropic describes trains that judgment on what its scientists chose to test rather than on what the bench later confirmed.

Somewhere inside a twenty-one hour search, one agent stopped and said this in its own working notes: "[The DNA next to the RT] is spectacular: I can see by eye a tandem repeat array … that's a CRISPR-like … repeat array?!" Anthropic quotes the line in the post it published on 23 September announcing a new life sciences group, a lab in the Bay Area, and a find it calls ART, for array-associated reverse transcriptases. The exclamation is the most revealing sentence in the announcement, and not for the reason it was included. The agent was not deriving new biology. It was recognising a shape it already knew, sitting somewhere nobody had reported one.

That is genuinely how this kind of discovery works, and it is worth saying plainly rather than deflating it. But the part of the campaign that did the work was not the noticing. Anthropic's own numbers make that clear: its agents gathered over 200,000 reverse transcriptases, picked out 3,500 candidate systems, and narrowed those to the twenty most compelling, which they wrote up as human-readable reports. One survived to the bench. The scarce resource in genome mining has moved from candidates to the judgment that throws candidates away, and the loop the company describes trains that judgment on what its scientists chose to test, not on what the bench later confirmed.

Start with what a reverse transcriptase is and why a row of repeats next to one is interesting. An RT copies RNA back into DNA, reversing the usual direction of transcription. Bacteria carry many of them, often as part of immune systems that fight off phages, the viruses that infect bacteria. Genome mining is the practice of finding such systems by reading sequence databases rather than culturing organisms: you take a family of genes, look for members nobody has characterised, and then look at what sits beside them. Neighbourhood is the signal. Genes that work together tend to be stored together, so an uncharacterised gene flanked by a suggestive partner is a better lead than the same gene alone.

The repeats are what made this one carry weight. A CRISPR array is a stretch of short, evenly spaced DNA sequences separated by spacers, and it functions as a bank of RNA templates. That bank is the reason CRISPR-Cas systems are programmable: change the stored sequence and you change what the enzyme is aimed at. So an array of evenly spaced repeats sitting next to an odd RT, plus an accessory protein of unknown function, suggests a system with the same architecture and therefore possibly the same trick. Anthropic says its first experiments show the ART array is indeed expressed as a set of distinct short RNAs. That is a real wet-lab result, not a pattern in a file.

Two things the announcement is careful about deserve the same care from anyone reading it. The RT itself, found in a jumbo phage, had already been identified in earlier studies; what Claude appears to be first to notice is the combination. And the primary function of ART is unknown. The system has been found, named and partly characterised. It has not been explained.

So look at the shape of the campaign rather than its result. Two hundred thousand to 3,500 to twenty to one is a generate-and-verify funnel, the same structure as best-of-n sampling with a scoring step, scaled up until the generator costs almost nothing. Roughly 950 agents and 210 million tokens sounds enormous, but it is the cheap half: at current prices it is a rounding error against a single set of bench experiments, and Anthropic frames it against weeks to months of expert time. What the funnel's quality depends on is not the 200,000 but the precision of the last cut. Twenty reports reached humans. Nineteen of them, on the company's own account of how a survey goes, were set aside.

Anthropic is admirably direct about what happens next, and this is the part that should hold a learner's attention: "Because Claude produces hypotheses so prolifically, the hypotheses themselves have become an object of study for us." With hundreds to thousands of candidate reports per campaign, the team says it has been asking what distinguishes the proposals it judges worth testing from those it sets aside, and that what it learns goes back into the instructions given to Claude, teaching it to mimic their scientific taste.

Read that as a supervised learning problem, because it is one, whether the result lands in a prompt or eventually in weights. The inputs are candidate reports. The labels are the team's decisions. And the labels record a preference, not an outcome. "We judged this worth testing" is a statement about the current state of expert intuition; "this turned out to do something" is a statement about the world. The second kind is the one that could correct a mistaken prior, and it is exactly the kind this loop cannot get yet. Bench work takes months, most candidates never reach a bench at all, and for ART itself the outcome label does not exist: nobody knows what the system does. A loop trained on shortlists learns to produce things that look like the shortlist.

That is not a hypothetical failure mode, it is the standard one. Learning a ranker from logged human choices biases it toward what was shown and picked; the candidates a reviewer never saw contribute nothing, and neither do the ones rejected wrongly, because a rejection produces no evidence about what it discarded. The asymmetry compounds when novelty is defined against an annotation. "Previously uncharacterised" means absent from what has been described in the literature at the time of the search. A model reading that literature is very good at spotting things that are anomalous with respect to it, which is precisely the CRISPR-shaped exclamation quoted above. It has no comparable handle on a system that resembles nothing described yet, because there is no known shape for it to be reminiscent of.

The strongest objection is that this is simply how science has always worked. Taste is transmitted by apprenticeship. Every generation's discoveries are shaped by the categories it inherited, and CRISPR was itself noticed as an odd repeat by people trained to find odd things. If prior-shaped selection invalidated ART, it would invalidate the whole discipline. There is a second, better version of the objection: a generator this cheap should let you lower the bar rather than raise it, because proposing costs nothing, so you can afford to test the strange candidate you would previously have skipped. And Anthropic did the expensive thing. It built a lab, ran expression and structural work, and published a result whose central fact is an absence of explanation.

All of that holds, and it still leaves the asymmetry. Generation scaled by three orders of magnitude in this campaign. Benches did not scale at all, and all the lab work is done by human scientists. When the ratio of proposals to experiments moves that far, the selection function is under more pressure than it has ever been, and its errors are the ones nobody can see. The number that would settle it is not in the post: across campaigns, how many candidates reached the bench, and how many became anything. One sentence in the announcement gestures at the denominator with real honesty, and it is the sentence I would build the evaluation around. A survey, it says, may end with a single candidate worth testing, or with none.

This matters beyond one company because the generator is being distributed and the verifier is not. Anthropic opened 10,000 subscription seats for scientists in August and offers up to $50,000 in credits per project, and its Life Sciences Verification Program, launched on 17 September, grants credentialed teams access to models with safeguards relaxed for biology. Hypothesis production is becoming abundant across the field. Wet labs, review capacity and funding are not. Whoever's taste decides which hypotheses get a bench becomes the effective author of the next decade of results, and that function is currently being distilled from a handful of expert shortlists.

For anyone learning this, three things transfer. Read "novel" as unannotated at search time, not as new in nature, and check what the reference database was. When you build a system like this, the artifact worth keeping is the rejection log, because the only way to measure a selection step is to go back and test some of what it discarded. And keep preference labels and outcome labels in separate columns, always: one tells you what an expert believed on a Tuesday, the other tells you what happened.

The pre-print behind this work sits on a domain this environment cannot reach, so everything above comes from Anthropic's own account, which is the only record of the campaign that exists. Take the numbers as the company's, not as independently checked.

There is a small piece of evidence in the naming. ART stands for array-associated reverse transcriptases: the system is named for what it sits next to. Restriction enzymes, Taq polymerase and CRISPR were named or repurposed once someone knew what they did. This one carries a location where a function should be, which is an accurate label for where the whole enterprise stands. The search got very cheap. Knowing what you found did not.

What this is argued from

Reporting and primary material the piece rests on, dated at the time of writing. The interpretation is mine; the facts belong to these.

  1. Claude discovers a novel enzyme system with CRISPR-like repeats Anthropic · 2026-09-23
  2. Introducing the Life Sciences Verification Program Anthropic · 2026-09-17
  3. Expanding our support for scientists Anthropic · 2026-08-27

Editorials on this site are written to be argued with. If you think the reading is wrong, it probably is in some particular way, and that is the useful part.

ai for sciencehypothesis generationagent searchselection biasgenome mining