Red-Team Evidence and Its Limits
What a red-team exercise contributes to an assurance case, why coverage cannot be quantified, and how to report results so they inform a decision rather than reassuring the reader.
Red-teaming has become a standard component of an AI assurance case and is frequently the weakest evidence in it, not because the work is poor but because of what the activity can establish. Understanding the shape of that limit determines how the results should be used.
What it contributes
Red-teaming is existence proof generation. Every finding is a demonstration that a failure is reachable, which is a strong and specific claim. It is the only method that reliably surfaces failure classes nobody anticipated, because it introduces adversarial creativity that no benchmark or automated suite contains.
It is also the activity that produces the regression cases and the concrete failure-mode documentation that other artefacts depend on, so its output feeds the behavioural test suite, the model card's limitations section, and the incident runbook.
What it cannot establish
Coverage. There is no way to quantify what fraction of the attack or failure space was explored, so no red-team result supports a statement about residual risk. "We found no further issues" is a statement about the team, the time and the imagination applied, not about the system.
Rate. A red-teamer eliciting a harmful output after forty attempts has established reachability. Whether that behaviour occurs once in forty ordinary interactions or once in forty million is a different measurement requiring representative traffic, and the two are routinely conflated in reporting.
Durability. A finding fixed by a targeted mitigation may be reachable by a nearby variation, so a closed finding is evidence about the specific instance and weak evidence about the class.
Reporting that informs
An assurance-grade red-team report states the scope tested and the scope excluded; the resources applied, in people, hours and expertise; the methodology, including whether attacks were adaptive; every finding with a reproduction and a severity based on harm rather than on cleverness; the disposition of each finding; and, explicitly, what was not covered.
The last item is the one that turns a reassuring document into a useful one, because a reader's question is what remains unknown and only the report can answer it.
When it breaks
Findings are graded on ingenuity rather than harm. An elaborate multi-step attack is more interesting and often less important than a trivial one that anyone could stumble into. Severity should track expected harm and the population exposed, and an interesting-findings ranking systematically misprioritises.
Internal teams share the builders' blind spots. A red team drawn from the same organisation, with the same assumptions about how the system will be used, misses the misuse patterns that come from unfamiliarity. External and community involvement is what widens it, and it is expensive.
Fixing findings can teach the wrong lesson. Patching each reported prompt individually produces a model resistant to those prompts and no more robust generally, and it produces a clean re-test that reads as improvement. Fixes should address the class, and the re-test should use held-out variations.
A red-team report is used as an assurance conclusion. Because it is concrete and readable, it gets cited as evidence the system is safe, which is the one thing it cannot show. Its place in an assurance case is as evidence of what was found and fixed, sitting alongside quantitative evaluation rather than substituting for it.
12 flashcards for this concept
Click a card to reveal the answer.