Evaluation & Evidence 2 October 2026 8 min read 1,662 words

Within one level is almost free

Anthropic says Claude now leads 26% of its own AI research, and validates that rating by reporting that its judge model and its staff agreed within one automation level 97% of the time. On a scale where nine tenths of the work sits on two adjacent rungs, that agreement is close to what guessing would give you.

The argument

The AL3-to-AL4 boundary that defines the index's 26% headline is the one distinction its own validation shows nobody can reliably make, because almost all the work sits on two adjacent rungs and agreeing within one rung therefore costs a rater nothing.

Anthropic's R&D Automation Index has six rungs and uses about two of them. More than nine tenths of the company's AI research and development work sits at AL3 or above, where "AI 'collaborates': it can do large chunks of work under close human direction". None of it sits at AL5, where a model would work with no human in the loop. The headline figure, that Claude "leads" 26% of Anthropic's AI R&D work as of August 2026, up from under 1% in February, is the share of that mass which cleared one rung rather than resting on the one below it.

This desk has already read the oversight numbers in the same report. The automation index is the first of its three measurements, and it is the one that will be quoted in arguments about how fast the frontier is moving, so it is worth reading at the same resolution. The most interesting paragraph in it is the validation: "Our judge model agreed with humans about as often as humans agreed with each other (model-versus-human exact agreement was 59%, human-versus-human was 35%), and model and human ratings were within one level of each other 97% of the time."

Read that against the shape of the scale and the reassurance inverts. When almost everything you are rating sits on two adjacent rungs, agreeing within one rung is nearly free. The AL3-to-AL4 boundary is the only cut that separates 26% from "above 90%", and it is precisely the distinction the validation shows the raters could not reliably make.

How the number is built

The construction is more careful than most self-reported metrics, which is why it repays scrutiny. For each week in July 2026, Anthropic randomly sampled 20% of staff from each department doing model R&D. Claude research agents read those people's weeks through Slack and internal documentation and extracted roughly 15,000 granular tasks. Claude then organised them into a hierarchy of "542 nodes at different depths, of which 378 are leaves", with leaves as specific as eval platform defect diagnosis and fixes, or serving incident postmortems.

Then the tree is fixed in place. "We freeze this tree so that every measurement we make happens against the same basket of work." For each leaf, a Claude agent researches how that kind of work is actually done across the company, restricted to evidence from the month being rated or earlier, and an independent Claude judge assigns an automation level on a scale Anthropic adopted from Epoch AI. Weighting is by person-time: each employee contributes one unit per week, split across the tasks they performed, and a category's weight is the sum of the person-time landing in it.

So the headline is a weighted share of a six-point ordinal scale, where nearly every rating was produced by the system being measured. The human ratings exist to check that, and staff rated their own work "without knowing what evidence the models had gathered or how they had judged that evidence". That blinding is the right design. The problem is in how the resulting agreement is interpreted.

What an agreement number costs

Raw agreement is the fraction of items two raters labelled identically, and on its own it means nothing, because it includes all the agreement you would get from raters who had never looked at the work. This is old ground in measurement, and the standard fix is to subtract the agreement chance alone would produce. Cohen's kappa, as scikit-learn's implementation states it, is \( \kappa = (p_o - p_e)/(1 - p_e) \), where \(p_o\) is observed agreement and \(p_e\) is the agreement expected if each rater assigned labels at random from their own marginal distribution. For ordinal scales the same library offers linear and quadratic weighting, so that disagreeing by one level is penalised less than disagreeing by three.

Put the index's own distribution into that formula. The report publishes two cuts rather than the distribution, so one level has to be inferred: above 90% at AL3 or better, 26% at AL4, nothing at AL5. That leaves roughly two thirds of the weight at AL3 alone, about a quarter at AL4, and under a tenth below AL3. Two raters drawing independently from those proportions would agree exactly about \(0.64^2 + 0.26^2 \approx 0.48\) of the time, near half. The humans agreed 35% of the time. Chance-corrected, that is a kappa of roughly minus a quarter: worse than independent guessing from the marginal.

The within-one-level figure fares worse. On this distribution, two raters can land more than one level apart only if one of them goes below AL3, and under a tenth of the weight does. Take the simplest reading, with that remainder sitting at AL2: independent draws then differ by more than one level about 5% of the time, so chance alone reaches roughly 95% within one level, two points under the report's 97%. The one thing that would widen that margin is a remainder spread further down the scale, and the report does not break it out.

One assumption is doing real work here, and it should be stated plainly rather than buried: the calculation treats the validation sample as having the same spread across levels as the person-time-weighted population, and the report does not publish the validation sample's own distribution. If staff were asked disproportionately about borderline tasks, chance agreement would be higher still and the margin thinner. If they were asked about an unusually spread sample, it would be lower. The direction of the argument does not depend on which: a raw agreement figure on a distribution this concentrated is not evidence of anything until it is compared against chance, and no such comparison appears.

What survives and what does not

The trend survives. Under 1% to 26% in six months is a move far too large for rater noise to manufacture, and that is the claim the index was built to support. But notice what the frozen basket does to it. The measurement asks how automated last July's work has become, and the report is candid that it "does not, on its own, tell us whether new kinds of work are appearing that humans have shifted onto". Person-time weights, which the authors call a crude approximation, point the same way: the more thoroughly a task is handed to a model, the less human time it consumes, and the smaller its share of the basket becomes.

What does not survive is the level. The 26% and the "above 90%" are not two findings. They are one distribution cut twice, and one of those cuts falls exactly where agreement is weakest.

There is a second reason that particular cut is soft, and it is in the definitions rather than the statistics. AL4 is "AI 'leads': it can complete most of the task end-to-end from a high-level prompt, while the human supervises." AL5 is illustrated by a case where "the engineer wouldn't even have to bring the issue to Claude's attention", and Claude "would be trusted to monitor for failures itself". Both hinge on what the human does, not on what the model can do. AL5 standing at zero is therefore a fact about Anthropic's supervision policy, truthfully measured, and a lab could raise its index by loosening that policy with no model changing at all.

The ambiguity of the rung below shows up in the lab's own published experience. In Claude-shaped science, the physicist Matthew Schwartz describes computing thirty elliptic Feynman integrals, half of them new, and producing 36 manuscripts across 18 fields with 19 coauthors in three months. He also reports that the results were often technically correct but scientifically unremarkable until expert collaborators redirected the work. A rater asked whether Claude led that work can say yes with a straight face. A rater who weights the redirection can say the human led. Both are reading the same episode against a one-sentence rubric, and this is the crack the 35% is measuring.

The strongest objection

Anthropic is, as far as I can establish, the only lab publishing anything of this kind, and it declares its own weaknesses: judge-model dependence, the frozen basket, person-time as a proxy, and an explicit call for third-party verification before anyone compares labs. A prototype index exists to establish a direction and invite others to publish theirs. Attacking the arithmetic of the only disclosure on the table is a good way to ensure the next lab publishes nothing.

That objection is right about the disclosure and wrong about the headline. The criticism here is not that the index was published; it is that the figure lifted out of it is the figure its validation supports least, and the fix is nearly free. Publish the full AL0-to-AL5 distribution, the validation sample's distribution across levels, and a weighted kappa. All three already exist inside the pipeline. The index would be harder to quote and much harder to dismiss.

What to take from it

Three habits generalise beyond this report. Never accept a raw agreement percentage without the marginal distribution, because the distribution sets how much of it was free. On an ordinal scale, ask for weighted kappa, since within-one-level claims on a concentrated scale are close to tautologies. And when the labeller is itself the system under study, remember what a single judge buys: variance collapses, so the number is highly reproducible, and bias goes entirely unmeasured. Reproducible is not the same as right.

The sharpest number in the report needs none of this machinery. AL5 is empty. Not because no model could close that loop, but because nobody inside Anthropic has yet been allowed to stop watching. That figure requires no rubric, no judge and no agreement statistic to read, and of everything the index reports it is the one that would actually change the day it moved.

What this is argued from

Reporting and primary material the piece rests on, dated at the time of writing. The interpretation is mine; the facts belong to these.

  1. Measurements for understanding the pace of AI development inside frontier labs Anthropic (Marina Favaro and Phillie Wright) · 2026-09-17
  2. When AI builds itself Anthropic (Marina Favaro and Jack Clark) · 2026-09-18
  3. Claude-shaped science Anthropic (Matthew Schwartz) · 2026-10-01
  4. cohen_kappa_score, sklearn/metrics/_classification.py scikit-learn on GitHub · 2026-10-02

Editorials on this site are written to be argued with. If you think the reading is wrong, it probably is in some particular way, and that is the useful part.

automation levelsinter-rater agreementllm-as-judgemeasurementrecursive self-improvement