Attacker Cost as the Robustness Metric
Why language-model defences cannot be certified and are therefore reported as attacker spend rather than as a bound, what evidence makes such a claim checkable, and the 2025 result that broke twelve defences which had each reported near-zero attack success.
Twelve recently published defences against jailbreaks and prompt injection, most of which had reported attack success rates near zero, were pushed above 90% ASR by attackers who simply spent more: tuned and scaled gradient descent, reinforcement learning, random search, and human-guided exploration aimed at each defence's specific design (Nasr et al., 2025, The Attacker Moves Second, arXiv:2510.09023). The authors are from OpenAI, Anthropic and Google DeepMind, which is itself part of the finding: this is not a dispute between labs about whose defence is better.
Vision had this moment twice. Athalye et al. found that 7 of the 9 non-certified white-box defences at ICLR 2018 relied on obfuscated gradients and broke 6 of them completely (Athalye et al., 2018, Obfuscated Gradients Give a False Sense of Security, ICML, arXiv:1802.00420). Tramèr et al. then broke thirteen more from ICLR, ICML and NeurIPS with adaptive attacks (Tramèr et al., 2020, On Adaptive Attacks to Adversarial Example Defenses, NeurIPS, arXiv:2002.08347). The difference is that vision had a fallback: a threat model precise enough to certify against.
Why there is no certificate for language
Certified robustness needs a formal set of admissible inputs. For an image classifier, \(\{x' : \|x' - x\|_\infty \le \varepsilon\}\) is a mathematical object, and randomised smoothing or interval bound propagation can prove no member of it changes the prediction, as Adversarial Training and Certified Defences sets out.
The language equivalent would be the set of all prompts that elicit a given harmful capability. That set has no metric, no projection operator and no closed form; it is defined by what a human judge would call the same request. Worse, the quantity to be bounded is not a label but a policy judgement about a free-text response. Nothing in the certification toolkit applies, and no amount of better optimisation will make it apply.
What remains is an economic claim: not "no attack exists" but "an attack costs more than it is worth". That is a weaker statement than a certificate and it is not a vacuous one. It is the form every claim in operational security takes.
Metrics that carry information
A cost-based claim is checkable only if the budget is reported with the rate. The quantities that do work:
| Quantity | What it measures | Example |
|---|---|---|
| ASR at stated \(N\) | Success under a fixed sampling budget | 78% at \(N = 10{,}000\) on Claude 3.5 Sonnet (Hughes et al., 2024) |
| Red-team hours to first universal jailbreak | Human effort, adaptive and creative | None in 3,000+ hours; one in about 3,700 more (Anthropic, 2025) |
| Detections per thousand queries | Residual risk at production scale | 0.005 per thousand over 198,000 attempts (Anthropic, 2026) |
| Cost to unalign | Spend to remove the defence outright | 10 examples, under $0.20 (Qi et al., 2024) |
| Defender overhead | What the defence costs to run | 23.7% compute in 2025, about 1% in 2026 (Cunningham, Wei et al., 2026) |
Two reporting rules make the difference between a measurement and a press release. State the attacker's budget alongside every rate, because a rate without a budget is unfalsifiable. And treat a defence's own evaluation as a lower bound on attack success, never an upper bound, since the authors optimised neither the attack nor against their own design.
When it breaks
Cost claims decay. Attack cost falls with every published technique, every cheaper model and every open-weight release that can serve as a surrogate. A robustness claim is dated in a way a certificate is not, so "robust as of early 2026" is the only honest tense.
Attacker cost is not uniform. A defence can be expensive to break at scale and cheap to break once. A universal jailbreak amortises across every query, which is why the universal case is tracked separately and why the per-prompt ASR is the less interesting number.
Absence of evidence is doing a lot of work. "No red-teamer found a universal jailbreak" is a statement about the red-teamers, their incentives and their time. It is real evidence, it is the best available, and it is not a bound; the limits of this reasoning are the subject of /learn/red-team-evidence-and-its-limits.
Empirical robustness still has a measurable ceiling in vision, and it is low. Adversarially trained CIFAR-10 models went from 45.8% robust accuracy at \(\varepsilon = 8/255\) (Madry et al., 2018, arXiv:1706.06083) to 70.69% with generated-data augmentation (Wang et al., 2023, ICML, arXiv:2302.04638), against 93.25% clean accuracy, after a decade of work on the easiest formalisable version of the problem. Anyone expecting the language version to be solved rather than priced should sit with that number.
References and further reading
Every source this page cites, in the order it cites them. All of them open in a new tab.
- Nasr et al., 2025, The Attacker Moves Second, arXiv:2510.09023 arxiv.org
- Athalye et al., 2018, Obfuscated Gradients Give a False Sense of Security, ICML, arXiv:1802.00420 arxiv.org
- Tramèr et al., 2020, On Adaptive Attacks to Adversarial Example Defenses, NeurIPS, arXiv:2002.08347 arxiv.org
- Hughes et al., 2024 arxiv.org
- Anthropic, 2025 anthropic.com
- Anthropic, 2026 anthropic.com
- Qi et al., 2024 arxiv.org
- Cunningham, Wei et al., 2026 arxiv.org
- Madry et al., 2018, arXiv:1706.06083 arxiv.org
- Wang et al., 2023, ICML, arXiv:2302.04638 arxiv.org
7 flashcards for this concept
Click a card to reveal the answer.