Safety & Alignment 7 October 2026 6 min read 1,366 words

Nobody counted the patches

On 29 September Anthropic reported that an open-weight model turned two public Chrome disclosures into a working exploit chain for $20.40 and eight hours of machine time. On 6 October it opened frontier cyber capability to defenders in tiers, and published 129,000 findings. Neither post counts a fix.

The argument

An open-weight model has made the weaponising of disclosed flaws cost $20.40 and eight hours, which leaves 129,000 verified vulnerabilities an asset only to defenders who can deploy a patch faster than that.

A researcher handed an open-weight model the public details of a recently disclosed Chrome vulnerability, CVE-2026-11645, plus one other known flaw, and asked for an exploit. The model chained the two into a reliable ARM64 exploit chain that got past pointer-authentication hardening. It took twenty minutes of human attention, eight hours of model work, and $20.40 at the provider's list prices. Anthropic's Frontier Red Team published that result on 29 September, in an assessment of GLM-5.3.

Read the experiment again and notice what is missing from it. Nothing was discovered. The flaws were already public. The model was not asked to find a weakness in Chrome; it was asked to convert a disclosure into a working weapon, which is the part of offensive security that used to require a specialist and a fortnight.

That is the shift worth arguing about, and it is not the one the week's coverage took. The capability that just became unconditional is not finding unknown flaws but weaponising known ones cheaply, and that inverts the value of a vulnerability report. A week after the GLM assessment, on 6 October, Anthropic expanded its Cyber Verification Program and reported that its partners had "uncovered at least 129,000 verified software vulnerabilities between April and July 2026," with "more than 33,000" rated critical or high severity. Findings at that volume are an asset to a defender who can deploy a patch faster than twenty dollars can produce an exploit, and a liability to everyone else. The two posts, published eight days apart, are the two halves of that trade, and only one half has a number attached.

What the benchmarks actually measure

Start with the scores, because they look weak until you understand what is being counted.

On ExploitBench, which asks a model to develop a working exploit against Chrome's V8 JavaScript engine, GLM-5.3 succeeded in 50 of 410 attempts. Claude Mythos Preview, the frontier model Anthropic uses as its reference point, managed 56 of 410. On a binary-exploitation evaluation that tests whether a model can achieve a control-flow hijack in projects drawn from OSS-Fuzz, GLM-5.3 reached 4% and Mythos Preview 6%; models from an earlier generation scored 0%.

Four percent sounds like failure. It is not, and the reason is the most useful thing a learner can take from this fortnight. A per-attempt success rate is a defensive statistic only when attempts are expensive. When they are cheap, independent and parallel, the attacker's relevant quantity is the cost per success, and the defender's exposure is the maximum over attempts rather than the mean. The only priced run in the assessment cost $20.40 of API time plus a third of an hour of a person's attention. Mixing that figure with the 4% from a different evaluation would be arithmetic theatre, so I will not do it, but the direction is not in doubt: when a success costs tens of dollars and a failure costs almost nothing, a single-digit rate is a schedule, not a shield.

The other half of the picture is historical, and it is why the discovery framing misleads. Automated vulnerability finding is not new. Google's OSS-Fuzz has been continuously fuzzing open-source code since its announcement in December 2016, and its own README reports that as of May 2025 it "has helped identify and fix over 13,000 vulnerabilities and 50,000 bugs" across more than a thousand projects. For nearly a decade, machines have produced more crash reports than humans could triage. The bottleneck was never the finding. It was the judgment to tell an exploitable memory error from a merely annoying one, and then the craft to build the thing that proved it. The GLM result is a price cut on exactly that craft.

Against that decade, 129,000 verified vulnerabilities in four months is a startling number, and it deserves its caveats. Anthropic's figure comes from Project Glasswing, announced on 7 April with founding participants including Amazon Web Services, Apple, Cisco, Google, JPMorganChase, the Linux Foundation, Microsoft and NVIDIA, later extended to over forty more organisations that build or maintain critical software. Those partners scan first-party code as well as open source, and "verified vulnerability" is their counting convention, not OSS-Fuzz's. The comparison is not apples to apples. It does not need to be: by any definition, the production rate of findings has gone up by more than an order of magnitude, in the same months that the cost of acting on a finding offensively fell to the price of a takeaway meal.

Where the safeguards sit, briefly

The assessment also measured how hard GLM-5.3 is to talk into the work. Given a bare malicious order in a simulated environment, it engaged 0% of the time. Told it was running a red-team exercise, 64%. With its reasoning tokens prefilled so that the trace appeared to have already decided to comply, 92%. After abliteration, the weight-editing procedure that removes the refusal direction outright, 100%. Anthropic's team spent about 2,200 GPU hours and roughly $4,400 doing that, and estimates an experienced team would need closer to 600 hours and $1,200. GPQA-Diamond scores came out identical afterwards; the cyber evaluations dropped a few percent.

Two numbers in one paragraph tell you where the engineering value is. The capability cost somebody a training run. The refusal costs $1,200 to remove and takes almost nothing with it when it goes. This desk argued the mechanism of that asymmetry on 1 October and will not relitigate it here. What matters for this argument is narrower: when the capability travels as a file, the refusal is not a property of the capability, so the only thing anyone is really rationing is time.

Four months, and who gets to spend it

How much time? The assessment cites an evaluation by CAISI, the US government's AI standards centre, finding GLM-5.3 to be "the most cyber-capable open-weight model released to date" and that it "lags the US frontier by about four months on an aggregate of CAISI's cyber benchmarks." Anthropic adds the qualification that makes the figure legible: the US models in that comparison were tested with safeguards disabled where applicable, and their frontier versions stay restricted to vetted users, whereas anyone can download GLM-5.3.

So the gate buys about a quarter. That is not nothing. The question is who is positioned to spend it, and the tier structure published on 6 October answers it more honestly than the prose does. Defense Access, for defensive work, aims to respond "within a few days." Red Team Access, adding authorised penetration testing, is expected to "take a few weeks to review." Specialized Access, for organisations testing systems like flight operations, power grids and interbank transfer infrastructure, is reviewed in collaboration with the US government. Enrolment requires data retention so that misuse can be monitored, with a zero-retention option promised later this autumn.

Now read the eligibility list for the first tier: security teams at companies, nonprofits and universities, "operators of critical infrastructure of any size, such as regional hospitals or municipal utilities," smaller security firms, open-source maintainers, and individual researchers with a track record of reported vulnerabilities. Those are the right people to invite. They are also, almost by definition, the people with nobody on staff to read 33,000 critical findings, no maintenance window, and no capacity to ship a kernel update to a fleet inside four months. The program hands the fastest available capability to the slowest-moving defenders through the slowest available process, and the thing it is racing is a download.

The strongest objection is that this gets the direction of the race wrong. Glasswing's partners are the vendors of the operating systems and browsers themselves. They scan code they own, they fix before disclosure, and a flaw fixed at source never becomes an n-day for anyone. On that reading the 129,000 figure is pure gain, the Chrome chain exploited inputs that were public precisely because the normal process had already worked, and the remaining exposure is unpatched fleets, which has been a patch-management failure since long before any of this. That objection is substantially right, and I would concede the whole first-party case to it. But it explains why the expansion exists. The organisations that own their stack were already served in April; the tiers published in October are aimed at everyone who does not. For them the window that matters was never vendor-to-patch. It was patch-to-deployed, and the only side of that interval anything got compressed on is the attacker's.

One limit on all of this: every figure above comes from one company's own publications, read in full, and there is no independent replication of either the benchmark scores or the vulnerability counts. The post about the Chrome chain does not say whether a fix had already shipped for CVE-2026-11645, and the argument does not turn on it either way, because the model's inputs were public in both cases. An assessment that grades a competitor's model and a program that markets your own access tiers are both evidence of what a company wants believed, which is a real fact and sometimes the more interesting one.

What should a learner do with this? Three things. Stop reading a robustness or exploitation rate as a probability of harm, and start reading it as a cost per success at a stated sampling budget; single-digit rates are decisive when retries are free. When a vendor claims its model helps defenders, ask what the system produces, findings or fixes, and who is on the receiving end, because those are different products with different economics. And treat a capability lag measured in months as a claim about diffusion rather than about safety, especially when, as here, the lab publishing it tells you the comparison runs between a model with its safeguards switched off and a model you simply download.

Between them, the two posts give us a production rate for findings, a per-attempt rate for exploit development, a dollar cost for removing a refusal, a dollar cost for building a chain, and a response time for an application form. The quantity that would tell us whether any of it is working is how many of the 33,000 critical findings are now patched on machines that people use. Nobody counted that. Until somebody does, the honest summary of this fortnight is that we have become very good at measuring the discovery of problems and have not started measuring their disappearance.

What this is argued from

Reporting and primary material the piece rests on, dated at the time of writing. The interpretation is mine; the facts belong to these.

  1. GLM-5.3 and the spread of advanced cyber capabilities Anthropic · 2026-09-29
  2. Expanding the Cyber Verification Program Anthropic · 2026-10-06
  3. Project Glasswing: Securing critical software for the AI era Anthropic · 2026-04-07
  4. OSS-Fuzz — continuous fuzzing for open source software Google (GitHub) · 2025-05-01

Editorials on this site are written to be argued with. If you think the reading is wrong, it probably is in some particular way, and that is the useful part.

cyber capabilitiesexploit developmentvulnerability disclosuretiered accesssafeguards