The limit was a price
A nine-loop scattering amplitude that had stood as a record since 2023 fell in late August to about a hundred dollars of processor time and a week of autonomous grinding. The interesting part is not what the model knew. It is what the field already had: a cheap way to prove it wrong.
The argumentClaude's ninth loop was never a harder calculation than the eighth, only an unpurchased one, and what this fortnight's two AI-science harnesses demonstrate is that a model's reach into a field is set by how cheaply that field can prove it wrong.
In early August a science writer who used to be a theoretical physicist put up a challenge: compute the nine-loop scattering amplitude in N=4 super Yang-Mills theory. The eight-loop result had stood since 2023, when Lance Dixon and Andy Liu reached it indirectly, through a related quantity called a form factor and a symmetry called antipodal duality. Before the month was out, Anthropic wrote to say Claude had the ninth loop, by two independent routes. The processor bill for the main calculation was about a hundred dollars, which Matt von Hippel describes as "running 96 CPUs for a week". Including the model's own time, he reckons either route would have cost an end user one or two thousand dollars.
Then comes the sentence that ought to unsettle anyone working in a quantitative field. "If people thought it was possible to just run the usual bootstrap method for one more loop," von Hippel writes, "someone would have done it."
Nobody thought it, so nobody spent the hundred dollars, and the record sat there for three years. That is the actual finding, and it is not a finding about intelligence. Claude's ninth loop was never a harder calculation than the eighth, only an unpurchased one. Put it beside the other AI-science harness published this fortnight and a sharper claim emerges: a model's reach into a field is set by how cheaply that field can prove it wrong.
Start with what a bootstrap is, because the whole argument lives in its shape. A scattering amplitude is the quantity that tells you how likely particles are to come in one way and leave another. You normally compute it as a series, and each additional loop is one more term, with combinatorially more ways for the particles to interact along the way. N=4 super Yang-Mills is nobody's description of the real world; physicists use it because its answers have unusually rigid structure, which makes it a proving ground for technique.
The bootstrap exploits exactly that rigidity. Rather than summing interactions, you assume the answer is built from a known vocabulary of functions, what von Hippel calls "a specialized alphabet", with unknown coefficients in front of each piece. Then, in his words, "you start checking everything you know: predictions from other calculation techniques, rules the answer has to obey, links to related problems where the answer was easier to find." Each thing you know becomes a condition the coefficients must satisfy. Impose enough conditions and only one candidate survives. His analogy is a good one: "a bit like Sudoku, where you begin with a grid with all possible numbers, then cross them out as you go."
Notice what kind of work that is. It is not derivation. It is propose-and-eliminate, and the elimination is mechanical. A wrong guess does not produce a plausible wrong answer; it produces a contradiction. Von Hippel calls the calculation "very fragile", meaning any slip kills it outright. Fragile and checkable are the same property seen from opposite sides. A procedure that collapses on a single error is a procedure that tells you, immediately and for free, that you have made one.
That is the property an agent can exploit, and it is a different property from being clever. The model does not have to be right. It has to be able to notice being wrong, thousands of times, at low cost, without losing patience or interest. Dixon, in an addendum to the piece, puts the field's side of this plainly: "We have all the tools to validate any candidate solution a machine would provide us." The amplitudes community had built its verification apparatus years before anyone pointed a model at it. The ninth loop was cheap because the marking was already free.
The second release this fortnight makes the point from the other direction. Matthew Schwartz, a Harvard physicist, published BootLoops, which he calls "a harness for large language models doing precision quantitative science", together with a report of what it produced: "36 manuscripts in 18 fields with 19 coauthors over three months, out of some 400 candidate problems." Read the repository rather than the summary and the engineering priorities are unmistakable. BootLoops 1.0 ships twelve skills, and seven of them are not research skills at all. They are discipline: an acceptance gate, planted-truth controls that "recover a known answer before any real data is touched", independence bookkeeping "so no oracle that fed a fit ever certifies the result", constant recognition with "refusal over invention when no relation is found", timing discipline, a reading contract. The toolkit's standard for a result is that it "reproduces an independent route at points no fit ever saw, with a positive control proving the check can fail".
A positive control proving the check can fail. That line is the whole discipline in nine words, and it is the opposite of what a year of product announcements has trained people to look for. The scarce component in both of these systems is not reasoning. It is a test that is capable of returning red.
Which is why the selection ratio matters more than the output count. Roughly four hundred candidate problems became thirty-six manuscripts. Schwartz is candid about the filter: the mismatch he set out to fix "is between what scientists want and what AI does well", and the problems that fit are quantitative ones where, as he puts it, "a technique from mathematics, physics, or computer science would solve outright if anyone knew it existed." Those are problems whose answers end in a number somebody can check. The harness is not a general accelerator for science. It is a sorting machine that finds the corners of science where verification is already cheap, and then buys the grinding.
So the honest version of "AI is accelerating research" is narrower and more useful than the headline: AI is currently repricing patience in fields that can mark their own homework. That has a consequence nobody is discussing. The visible record of AI-assisted science will be drawn disproportionately from verification-rich areas, because those are the only places the loop closes. Exact calculation, formal proof, enumeration with completeness certificates, Bayesian evidence to a proven error radius. Fields where the ground truth arrives in six months from a wet lab, or never, will appear to be falling behind when in fact they are merely unmarkable. Treating the first group's progress as a measure of model capability, and the second group's silence as a measure of its own, is the mistake this evidence invites.
The strongest objection is Dixon's own, and it deserves stating at full force. Writing fragile algebra that works first time, over a week, without a human debugging it, is not nothing. A graduate student with a hundred dollars of cluster time could have attempted this at any point in the last decade and would most likely have produced a mess. Dixon notes that while his team had the tools to check any candidate, the direct bootstrap route looked prohibitive enough that he expected the next loop to come indirectly, "potentially by a different kind of AI method". The capability is real.
Granted, and it is worth being precise about what the capability is for. It is sustained, error-free execution of a procedure that announces its own failures. That is an enormous thing to have acquired, and its reach is still bounded by the availability of the announcement. Two further details sharpen this rather than soften it. Von Hippel reports that Song He's group at the Chinese Academy of Sciences had already obtained most of the same result with GPT-6 assistance, arriving within days. When two different systems land the same answer in the same fortnight, that is evidence about the problem's ripeness, not about either model. And von Hippel's own conclusion is not triumphal. "I wanted to know how far AI could push computational limits," he writes, "and I feel like what I learned here is just that I was too naïve about where the limit was."
For anyone learning this field, that suggests three habits worth forming. When a result is announced, ask what graded it and what the grading cost, before asking what model produced it; the answer to the second question explains much less than it appears to. Read "AI solved X" as a statement about the verification economics of X rather than about intelligence, and notice how often that reading survives contact with the method section. And if you want to work usefully with these systems, the useful work is not in the prompt. It is in building the check that can fail, with a positive control to prove it can, which is ordinary scientific hygiene that models now make expensive to skip.
One caveat on the evidence: both accounts here come from the people who did the work, and the underlying papers sit on preprint servers this environment could not open. Dixon's validation is independent, and the BootLoops repository is inspectable, but neither the 2023 eight-loop paper nor Song He's result was read directly.
There is a last discomfort in the nine-loop story that has nothing to do with Claude. For three years, a result that several groups wanted was sitting behind a hundred dollars and a tolerance for tedium, and the reason nobody collected it was that everyone had concluded it was out of reach. The model did not move the limit. It revealed that the limit had been a price all along, and the only reason anyone knows that now is that something was finally willing to pay it. The question this leaves is not what else AI can do. It is how much of your own field is priced like that, and how you would ever find out.
What this is argued from
Reporting and primary material the piece rests on, dated at the time of writing. The interpretation is mine; the facts belong to these.
Editorials on this site are written to be argued with. If you think the reading is wrong, it probably is in some particular way, and that is the useful part.