Evaluation & Evidence 11 October 2026 7 min read 1,618 words

Eleven of the twelve were real

Anthropic's free vulnerability scanner for open source now mails model-generated reports to maintainers with no human review, and its own validation shows the dominant error is not invention but repetition. A fuzzer's crash carries a signature. A report written in prose carries nothing a tracker can key on.

The argument

Machine-scale vulnerability discovery is limited by identity rather than accuracy, because a fuzzer's crash arrives with a signature and a model's finding arrives as prose, so the same bug reported twice is counted as two true positives and nothing in the delivery path can notice.

To validate an early version of OSS Scanner, Anthropic gave the penetration testers who normally review its vulnerability disclosures a batch of 97 critical and high-severity findings drawn from 48 projects. Eighty-five of them, 88%, met the bar for the company's coordinated disclosure process. Which leaves twelve, and the useful part of the exercise is what was wrong with those twelve. One was a false positive. The other eleven, in the Frontier Red Team's own words, "were real but duplicated known issues or other findings from the scan."

Read that again, because it inverts the thing everybody expects from a language model pointed at a codebase. The failures were not inventions. They were repetitions. Only one finding in 97 described a vulnerability that did not exist; eleven described one that did, and had already been described, either by the project's own records or by another finding in the same batch.

That is the shape of the problem now, and it is not the shape the announcement is measured in. The Cyber Mission post says of the service, "We expect a true-positive rate above 90%, and will work to improve the true positive rate and fix quality over time." A duplicate is a true positive. The headline quality figure for OSS Scanner is calculated in a unit that cannot see the error that dominated its own validation. What limits machine-scale vulnerability discovery is not whether a finding is real. It is whether anyone can tell that two findings are the same finding.

What a scan actually hands over

OSS Scanner, announced on 8 October, gives opted-in open-source projects "thorough, periodic security scans by our strongest models at no cost". Enrolment is a pull request to anthropics/oss-scanner adding one directory with a project.yaml. Two fields are required: the repository to scan, and a primary_contact email. You must also supply a Dockerfile that installs every dependency and builds the project, because the scanner builds it online in an isolated virtual machine and then moves it to a network with no internet access before the audit begins. The reports, Anthropic says, are "fully model-generated, without human review or triage"; each contains a self-contained reproducer, an explanation including a bisection to the commit that introduced the bug where possible, and a candidate patch when one is available. They are emailed to primary_contact, they are not made public, and they carry no ninety-day disclosure deadline.

Now compare the thing it is explicitly modelled on. Google's OSS-Fuzz has pointed automated discovery at open source for a decade, and the machinery underneath it, ClusterFuzz, defines a term that OSS Scanner has no equivalent of. A crash state is, per the glossary, "a signature that we generate from the crash stacktrace for deduplication purposes". That one line is load-bearing for everything else in the pipeline. Because a crash has a signature, a crash has a row in a tracker. Because it has a row, the same crash arriving a second time is recognised and discarded rather than mailed. Because it has a row, the fix can be checked by re-running the reproducer and the row closed without a human deciding anything. And because the row has a creation date, the ninety-day clock in OSS-Fuzz's disclosure guidelines has something to attach to.

None of that follows from a well-written paragraph. The identity of a bug is not in the text describing it. The same off-by-one can be reached through three call paths and will be written up three ways; two findings can share a root cause and describe different symptoms; a sampler run a fortnight later on the same source tree will produce prose that is nowhere near token-identical to last fortnight's prose about the same defect. Exact hashing is useless here, and the fuzzy methods the field uses elsewhere for near-duplicate detection are tuned to catch reworded copies of the same document, not two different observations of one fault. A reproducer would be the natural key, except that a reproducer for a logic flaw is a program, and program equivalence is undecidable in general and expensive in practice.

So the receiving end of this service has no deduplication primitive at all. It has an inbox.

Anthropic is unusually candid about why the service exists in this form. Over the past six months its models produced over 29,000 candidate vulnerabilities in important open-source software, of which roughly 6,000 have been manually reviewed and triaged. "We remain bottlenecked on our human capacity to validate these findings," the post says. Nearly 5,000 reports have gone to maintainers already, some of them unvalidated, because maintainers asked for everything. The capability curve behind the volume is steep: on CyberGym, an academic vulnerability-finding benchmark, the post reports models going from under 20% of vulnerabilities at the start of last year to over 85% this year.

OSS Scanner's answer to a human-review bottleneck is to delete the human review. That is a coherent engineering decision, and it is honest about its cost: faster and more frequent scanning, with the acknowledgement that "some will contain inaccuracies, such as a wrong severity rating". But deleting a filter does not create capacity. It relocates the filter to the only place left, which is a maintainer reading email, and that place is the one with no signature, no tracker, no automatic fix verification and no state carried from one scan to the next.

Look at what the system does offer a maintainer who wants to stop hearing about something. project.yaml has a disabled flag, which pauses all reports. And there is threat_model.md, optional but strongly recommended, whose template asks the project to write down where untrusted input enters, which components are in and out of scope, how it rates severity, and — a section headed "Anything to leave alone" — with the worked example "do not report unaligned access in the SIMD code paths on x86".

That is a suppression list, written in prose, read by a sampler, with no mechanism that reports whether it was honoured. It is the only knob between "everything" and "nothing", and maintaining it is now part of the job. The work has not been removed from the maintainer. It has been changed from finding bugs into writing and revising a natural-language specification of which true findings not to send.

The case for the firehose

The strongest objection comes from the maintainers, not from theory, and it deserves stating at full strength. They are asking for more of this, not less. Todd Ouska of wolfSSL, quoted in the announcement: "We found the signal from these reports high: of the 74 reports we received, all but two were valid, and five became CVEs." Anthropic says maintainers receiving verified reports simply ask for bulk delivery of the unverified ones. A duplicate costs a few minutes of recognition. A missed remote-code-execution chain costs considerably more, and the asymmetry is not close. Nor is OSS-Fuzz's dedup free of its own failures: crash-state signatures split one bug across several rows and merge genuinely distinct bugs into one, as anyone who has worked a fuzzing queue knows.

All true, and it holds at today's volume. Thirty-three projects had been merged into the enrolment repository by 11 October, among them OpenSSL, NSS, Botan and cryptsetup, which are projects of exactly the kind that already run a security process. wolfSSL triaged 74 reports because wolfSSL is a company with people whose job that is. The design question is not whether a high-signal batch is welcome once. It is what the same inbox looks like under the stated goal of "faster and more frequent scanning", on the twentieth periodic scan, for a project whose maintainer has a day job. And notice the arithmetic buried in that friendly quote: 72 valid reports, five CVEs. Sixty-seven reports were real and did not become a tracked vulnerability. Some of those will be in-scope bugs fixed quietly, which is a fine outcome. Some will be the eleven-of-twelve case, arriving again.

The lesson generalises past security, and it is worth carrying into any system where a deterministic detector is being swapped for a model. The first property you lose is not accuracy. It is the identifier. A static analyser emits a rule ID, a file and a line; a fuzzer emits a stack signature; a failing test emits its own name. Each of those is a key, and a key is what lets a stream of findings be deduplicated, suppressed, tracked, aged and closed. Model output arrives as an argument instead, and an argument has no primary key.

So when you are offered a precision figure for a generated finding stream — security, code review, data quality, moderation — it is the wrong measurement to accept. Ask instead for distinct actionable findings per hour of reviewer time, and ask what happens on the second run over unchanged code. If the answer involves a human remembering, there is no pipeline there yet, only a demonstration. And if you are on the receiving end, build the key yourself before the volume arrives: normalise each report to something stable, a tuple of file, function and bug class, and match against it, knowing the match will be imperfect. An imperfect key beats an inbox.

Anthropic forecasts that in two years AI will favour defence, and concedes the near term may not. The eleven duplicates suggest where the hinge actually is. The models can already find the bugs. What nobody has built is the thing that notices they have found one before.

What this is argued from

Reporting and primary material the piece rests on, dated at the time of writing. The interpretation is mine; the facts belong to these.

  1. An opt-in vulnerability-finding service for open-source software Anthropic Frontier Red Team · 2026-10-08
  2. OSS Scanner enrolment repository Anthropic · 2026-10-08
  3. Introducing the Anthropic Cyber Mission Anthropic · 2026-10-08
  4. OSS Scanner project.yaml template Anthropic · 2026-10-08
  5. OSS Scanner threat_model.md template Anthropic · 2026-10-08
  6. OSS Scanner contributing guide Anthropic · 2026-10-08
  7. ClusterFuzz glossary Google ClusterFuzz · 2026-10-11
  8. OSS-Fuzz bug disclosure guidelines Google OSS-Fuzz · 2026-10-11

Editorials on this site are written to be argued with. If you think the reading is wrong, it probably is in some particular way, and that is the useful part.

vulnerability discoverydeduplicationtriagefalse positivesopen source