Reliability & Consequence 9 October 2026 7 min read 1,628 words

The clock did not scale

Automated vulnerability discovery has arrived at industrial volume, and the ninety-day disclosure deadline has been dropped to let it in. The instrument that used to put a date on someone else's fix is gone, and nothing outside your own organisation will ever date one again.

The argument

Model-generated vulnerability reports cannot carry a disclosure deadline, because nobody verified them — so the only mechanism that ever converted a finding into a scheduled fix has been retired exactly as the volume of findings made it indispensable.

The repository went up on 8 October with a single commit. A day later it had fifty-five open pull requests and no projects/ directory, because nothing had been merged yet. Each of those pull requests is a maintainer of some piece of widely used open-source software volunteering their project for a free, periodic security audit by a frontier model — asking, in effect, to be told what is wrong with their code. They are queuing. Admission is case by case, restricted to core maintainers, and gated on a contributor licence agreement.

The interesting sentence is not in the announcement. It is in the enrolment README, explaining why these audits work differently from the ones everybody already knows: "Because of this we do not place a 90-day disclosure period on these findings and will not make them public."

That is a small sentence about a large retirement. OSS Scanner is explicitly modelled on Google's OSS-Fuzz, which has spent years pointing automated discovery at open source under the ninety-day rule it takes from Google's own disclosure policy: ninety days after notification, or the release of the fix, whichever comes first, with a fourteen-day grace period if a patch has a date. We have always discussed that clock as a disclosure policy, an argument about transparency and the public's right to know. It was never mainly that. The clock was a scheduling instrument. It took a defect sitting in someone else's backlog, a backlog over which no outsider has any authority, and gave it a date. It is the only mechanism the industry has ever built that reliably converted a finding into a shipped fix, and it worked by making the fix someone else's deadline.

It has now been dropped at the precise moment the volume of findings made it matter most.

The volume is not speculative. Anthropic's Frontier Red Team says it has found "over 29,000 candidate vulnerabilities" in important open-source software in the last six months, and has "only been able to manually review and triage approximately 6,000 of these" — a human triage capacity running at roughly a fifth of machine discovery, which the team names plainly as the bottleneck. Two days earlier, the company reported that partners in its Project Glasswing programme "uncovered at least 129,000 verified software vulnerabilities between April and July 2026", with more than 33,000 of those rated critical or high. That figure comes from thirty-three partner reports and partial survey data, and the company expects "the true impact to be at least five times higher". The capability curve behind it is measurable: on CyberGym, an academic vulnerability-finding benchmark, the post reports models going from under 20% of vulnerabilities at the start of last year to over 85% this year.

Set against those numbers, one figure is missing, and the vendor says so: "fewer than 50% of partners disclosed patched numbers, often because their fixes were still in progress, so the patch rate is significantly undercounted." We know what was found to four significant figures. Nobody knows what was fixed.

The honest reading is not that someone forgot to instrument remediation. It is that the artefact changed character, and the clock could not survive the change. A fuzzer crash is self-verifying: the reproducer either faults the binary or it does not, and the same infrastructure that found it can confirm the fix landed. That property is what made a deadline defensible. A model's report is a different object. It carries a reproducer and often a candidate patch, but also a severity rating and an implicit judgement about the project's threat model, and those are opinions. In validation, expert testers checked 97 critical and high findings across 48 projects; 85 met the bar for coordinated disclosure, eleven of the remaining twelve were real but duplicated other findings, and one was a false positive. The company expects a true-positive rate above 90% on unreviewed output. Attaching a publication deadline to a stream like that would mean threatening to expose maintainers for failing to service a queue they did not ask for, on evidence that is wrong often enough to matter. Nobody should do that. So the clock goes — not out of timidity, but because an unverified finding has nothing a deadline can bite on.

What takes its place is capacity-matching, and it is worth looking at squarely, because it inverts something architects are used to. OSS Scanner, the Cyber Mission post says, "is meant for projects with the capacity to keep up with surfaced findings"; projects without that capacity continue to get slower, human-verified disclosures. Anthropic has funded two intermediaries, Akrites and Gold Eagle, specifically to collect and coordinate reports "to avoid overwhelming maintainers". All of this is sensible, even admirable. It also means that how much a project is told about its own defects is now a function of how much it can repair. Visibility has become downstream of throughput. The eligibility review, the CLA, the queue of fifty-five — that is a system rationing knowledge by the recipient's ability to act on it, because the alternative is dumping findings on volunteers and calling it help.

The strongest case against treating any of this as a problem comes from the maintainers themselves, and it is not a straw man. They are asking for more, not less. Anthropic reports that maintainers receiving its first verified reports "simply ask us for a bulk submission of all the unverified reports with proposed patches", and that nearly 5,000 raw reports have gone out on that basis. wolfSSL says that of 74 reports received, all but two were valid and five became CVEs, and that "with patches attached, the reports slotted right into our existing process". OpenSSL's Anton Arapov says the reports, "raw model output included, were as good and sometimes better than what we get from people". PostgreSQL's Noah Misch says fast-track access let the project address issues before they reached a general-availability release. If the people on the receiving end want the firehose, who exactly is harmed by removing a deadline nobody wants?

Notice which projects those are. OpenSSL, PostgreSQL, wolfSSL — organisations with a security process, a release train, paid staff and a professional relationship with their own backlog. They are the ones for whom the clock was always redundant, because they had their own. The ninety-day rule never existed to discipline them. It existed for everyone else: the vendor whose patch requires a customer maintenance window, the embedded supplier whose product shipped three years ago, the enterprise where upgrading one library is a quarter's work. Those organisations did not opt into anything this week, and they are the ones who will now receive findings — through their suppliers, through scanners built on the same models, through the security vendors in the new critical-infrastructure programme — with no external date attached to any of them.

Anthropic is clear-eyed about where that leaves the near term. "The cost of exploiting vulnerabilities has dropped," the Cyber Mission post says, "while verifying, disclosing, and fixing them is slow and still depends on people." In Glasswing, months routinely passed between a vulnerability being found and being fixed. For operational technology, where a fix waits for a window in which it can be applied safely to running machinery, the post says that in rare cases "this might take decades". The company forecasts that AI will favour defence in two years, and concedes that in the near term it may not. An attacker, meanwhile, submits no pull request, signs no CLA, and passes no eligibility review.

So the architectural work is not better scanning. It is manufacturing the clock yourself, as a property of the system rather than a policy on a wiki. The question that now determines a security posture is not how many open findings you have — that number has become a statement about how much scanning you bought — but how long it takes you to get a patch from report to production, and whether anyone owns that figure. Which components can be rebuilt and redeployed without a human deciding to? Which dependency bump is a ticket, and which is a project? Where in the estate does a fix require a customer's permission, and what is the longest such path? A programme that produces 129,000 findings in four months is answering an engineering question for you: the constraint was never knowing.

The one project operating at industrial finding volume for years has already written the answer down, and it reads differently now. The Linux kernel assigns CVE numbers to "any bugfix that they identify", refuses to assign them for issues that are not yet fixed, declines to tell you which ones apply to your system, and advises taking all released changes rather than cherry-picking individual fixes. Read in 2024 that sounded like a maintainer's convenience. Read against a scanner that can generate findings faster than anyone can triage them, it is the only posture that survives: make the fix the unit, and make taking fixes continuous and unremarkable, because the findings have stopped being individually decidable.

OSS-Fuzz still states its own record as a single number — it had, as of May 2025, "helped identify and fix over 13,000 vulnerabilities" — because its pipeline closed the loop, and the deadline was what closed it. The new programmes can tell you 29,000 found and 6,000 triaged. The third number does not exist. For ten years the ninety-day clock meant that at least one date in the repair process was not yours to negotiate. From now on, every date is. Most organisations have never started that clock, and nothing arriving in their inbox will start it for them.

What this is argued from

Reporting and primary material the piece rests on, dated at the time of writing. The interpretation is mine; the facts belong to these.

  1. Expanding the Cyber Verification Program Anthropic · 2026-10-06
  2. Introducing the Anthropic Cyber Mission Anthropic · 2026-10-08
  3. Launching an opt-in vulnerability-finding service for open-source software Anthropic Frontier Red Team · 2026-10-08
  4. OSS Scanner enrolment repository Anthropic · 2026-10-08
  5. OSS-Fuzz bug disclosure guidelines Google OSS-Fuzz · 2026-05-07
  6. OSS-Fuzz README Google OSS-Fuzz · 2026-01-30
  7. CVEs — Linux kernel documentation Linux kernel · 2024-02-17

Editorials on this site are written to be argued with. If you think the reading is wrong, it probably is in some particular way, and that is the useful part.

vulnerability disclosureremediationpatchingopen sourceaccountability