Adversarial Robustness Without a Norm Ball: What Jailbreaks Did to a Decade of Theory
Twelve published defences against LLM jailbreaks, most reporting attack success rates near zero, were pushed above 90% in 2025 by attackers who mainly spent more compute. Adversarial robustness for language models has no certificate, no epsilon and no projection operator. What it has instead is a price, and the field is still learning to quote it honestly.
In October 2025 a team drawn from OpenAI, Anthropic and Google DeepMind took twelve recently published defences against jailbreaks and prompt injection, most of which had reported attack success rates near zero, and pushed almost all of them above 90% (Nasr et al., 2025, The Attacker Moves Second, arXiv:2510.09023). They did not discover a new class of attack. They took known techniques, gradient descent, reinforcement learning, random search and human-guided exploration, tuned each one against the specific design of each defence, and spent real compute.
Three labs that compete on safety claims published that result together. That is worth pausing on, because the finding is not about any one defence. It is about what the published numbers were measuring.
Image classifiers went through this twice. In 2018 Athalye and colleagues found that 7 of the 9 non-certified white-box defences accepted at ICLR that year relied on obfuscated gradients, and broke 6 of them completely (Athalye et al., 2018, Obfuscated Gradients Give a False Sense of Security, ICML, arXiv:1802.00420). Two years later Tramèr and colleagues broke thirteen more from ICLR, ICML and NeurIPS (Tramèr et al., 2020, On Adaptive Attacks to Adversarial Example Defenses, NeurIPS, arXiv:2002.08347). Vision survived that reckoning because it had somewhere to retreat to: a threat model precise enough to prove things about. Language does not have one, and no amount of better optimisation will produce one.
Why this matters: Jailbreak robustness cannot be certified, because the set of prompts that elicit a capability has no metric, no projection operator and no closed form. Every defence that has survived contact with adaptive attackers works by raising the cost of search. That makes robustness an economic claim, which is a weaker thing than a bound and a much more useful thing than a vibe, provided you quote the attacker's budget with every number you publish.
TL;DR
- Attack success against a language model is a function of attacker budget, not a property of the model. Best-of-N jailbreaking, which applies random capitalisation and shuffling with no optimisation at all, reached 78% attack success on Claude 3.5 Sonnet at 10,000 samples and at least 52% on every LLM tested (Hughes et al., 2024, arXiv:2412.03556).
- That budget curve is predictable: fitting on 1,000 samples forecast the rate at 10,000 with a mean error of 4.6 percentage points.
- Prefilling, high-temperature sampling, adversarial suffixes and fine-tuning attacks are one vulnerability seen from four interfaces: safety alignment lives in the first few output tokens (Qi et al., 2024, arXiv:2406.05946). Ten examples and under $0.20 of fine-tuning removed it from GPT-3.5 Turbo.
- A vision or audio encoder hands back the continuous input space: a 336 by 336 image patch gives an attacker roughly 339,000 differentiable parameters against the 20 tokens a GCG suffix optimises.
- Classifier layers work, and they are priced in three currencies, not one: constitutional classifiers cut automated jailbreak success from 86% to 4.4% for a 0.38 percentage-point rise in production refusals and 23.7% extra compute (Sharma et al., 2025, arXiv:2501.18837). By 2026 a two-stage cascade cut the compute to about 1% and refusals to 0.05%.
- The same classifier system survived more than 3,000 red-team hours with no universal jailbreak, then produced one under about 3,700 more hours from a larger crowd. The system did not change. The budget did.
- Vision's certifiable version of this problem is not solved either. A decade of adversarial training moved CIFAR-10 robust accuracy at \(\varepsilon = 8/255\) from 45.8% to 70.69%, against 93.25% clean accuracy. Expect language robustness to be priced, not solved.
At a Glance
flowchart LR
A["Attacker budget: samples, hours, dollars"] --> B["Attack surface: text, prefill, weights, image, audio"]
B --> C["Defence layers: alignment, classifiers, representations"]
C --> D["Outcome: success rate at that budget"]
D -->|"no certificate exists"| E["Robustness quoted as a price"]
E -->|"attack cost falls monthly"| A
classDef blue fill:#1e40af,stroke:#3b82f6,stroke-width:1px,color:#fff
classDef purple fill:#6d28d9,stroke:#a78bfa,stroke-width:1px,color:#fff
classDef teal fill:#0e7490,stroke:#22d3ee,stroke-width:1px,color:#fff
classDef amber fill:#b45309,stroke:#fbbf24,stroke-width:1px,color:#fff
classDef rose fill:#be123c,stroke:#fb7185,stroke-width:1px,color:#fff
class A blue
class B rose
class C purple
class D teal
class E amberThe loop at the bottom is the part that has no analogue in the certified world. A proof about an \(\varepsilon\)-ball does not expire. A price does.
Before the Norm Ball Broke
Adversarial examples started as a geometry problem. Szegedy and colleagues found that imperceptibly perturbed images flip confident classifiers, and the perturbation was small in a precise sense: small in \(\ell_\infty\), or \(\ell_2\), or whichever norm you chose to formalise "imperceptible" (Szegedy et al., 2014, Intriguing Properties of Neural Networks, ICLR, arXiv:1312.6199).
That choice of norm did enormous work. It made the threat model a mathematical object, \(\{x' : \|x' - x\|_\infty \le \varepsilon\}\), with a projection operator you could run inside an optimisation loop. Projected gradient descent followed, and with it the min-max formulation that remains the only empirical defence to survive adaptive attack:
Madry and colleagues trained against the inner maximisation and reported 45.8% accuracy on CIFAR-10 against white-box PGD at \(\varepsilon = 8/255\), with 87.3% clean accuracy (Madry et al., 2018, Towards Deep Learning Models Resistant to Adversarial Attacks, ICLR, arXiv:1706.06083). Low, honest, and reproducible. Certified methods arrived alongside, proving that no perturbation inside the ball changes the prediction, at the cost of looser guarantees and smaller radii.
Then came language models, and with them a problem that looked similar and was not.
timeline
title From epsilon-balls to attacker budgets
2014 : Szegedy et al. formalise adversarial examples inside a norm ball
2018 : Madry et al. make PGD adversarial training the standard
: Athalye et al. break 6 of 9 ICLR defences via obfuscated gradients
2020 : Tramer et al. break thirteen more defences with adaptive attacks
2023 : Zou et al. carry gradient attacks to text with GCG suffixes
: Carlini et al. jailbreak multimodal models with adversarial images
2024 : Qi et al. show safety alignment is only a few tokens deep
: Hughes et al. fit power laws to best-of-N attack success
2025 : Sharma et al. ship constitutional classifiers to production
: Nasr et al. push twelve LLM defences above ninety percent
2026 : Cunningham et al. cut safeguard compute roughly fortyfold with cascades[IMAGE: Two side-by-side panels. Left panel: a 2D point with a dashed square around it labelled "epsilon-ball, projection defined, certification possible". Right panel: the sentence "how do I synthesise X" surrounded by a scattered cloud of unconnected paraphrases, role-plays, encodings and faux dialogues, with no boundary drawn. Caption: "The threat model did not get larger. It stopped being a set."]
What Replaced the Threat Model
The certification gap is structural, not temporary
To certify robustness you need a formal set of admissible inputs and a decision to protect. The language version would be the set of all prompts eliciting a given harmful capability. Write it down and three things fail at once.
There is no metric: the distance between "explain the synthesis route" and a Base64-encoded role-play eliciting the same content is a judgement about meaning, not a number. There is no projection, so even granted a metric you cannot project an arbitrary point back into the set of valid natural-language prompts, which every certification algorithm needs in its inner loop. And the protected quantity is a policy judgement about free text, not an \(\arg\max\) over classes.
None of those is a gap better mathematics will close. They follow from the problem being specified in natural language. What is left is an operational-security claim: not "no attack exists" but "an attack costs more than it is worth".
Alignment turned out to be a prefix policy
The sharpest mechanistic result in this area explains four unrelated-looking attacks with one observation. Safety alignment adapts a model's generative distribution mostly over its first few output tokens.
Write \(p_\theta(y \mid x)\) for the aligned model. Alignment training moves \(p_\theta(y_{1:k} \mid x)\) substantially for small \(k\), while leaving \(p_\theta(y_{k+1:T} \mid x, y_{1:k})\) close to the unaligned base model's conditional (Qi et al., 2024, Safety Alignment Should Be Made More Than Just a Few Tokens Deep, arXiv:2406.05946). Condition the model on "Sure, here is a" and the refusal behaviour is mostly gone, because almost no alignment gradient ever reached that region of the sequence.
Every interface that lets an attacker influence those first tokens is therefore an attack.
flowchart TB
subgraph surfaces["Four interfaces and one vulnerability"]
P["Prefill the assistant turn"]
T["Raise temperature or top-k"]
S["Optimise an affirmative suffix"]
F["Fine-tune on compliant openings"]
end
P --> K["Attacker controls first k tokens"]
T --> K
S --> K
F --> K
K --> B["Later tokens follow base-model conditional"]
B --> H["Harmful completion"]
classDef purple fill:#6d28d9,stroke:#a78bfa,stroke-width:1px,color:#fff
classDef amber fill:#b45309,stroke:#fbbf24,stroke-width:1px,color:#fff
classDef rose fill:#be123c,stroke:#fb7185,stroke-width:1px,color:#fff
class P,T,S,F,H rose
class K amber
class B purple[IMAGE: Token-position heatmap of a 200-token assistant response. Positions 1 to 6 are deep red, labelled "alignment gradient concentrated here"; positions 7 onward fade to near-white, labelled "base-model conditional, largely untouched". Four arrows labelled prefill, temperature, suffix and fine-tune all point into the red band. Caption: "Shallow safety alignment, and the four interfaces that aim at the same six tokens."]
This reframes GCG. The attack optimises a suffix \(s\) to minimise \(-\log p_\theta(y^\star_{1:k} \mid x \oplus s)\) for a short affirmative target \(y^\star\), which always looked like an odd objective: why optimise for "Sure, here is" rather than for the harmful content itself? Because under shallow alignment the prefix is the whole fight. Zou and colleagues reported transfer rates of 86.6% to GPT-3.5, 66.0% to PaLM-2, 47.9% to Claude-1 and 46.9% to GPT-4, against 2.1% to Claude-2 (Zou et al., 2023, Universal and Transferable Adversarial Attacks on Aligned Language Models, arXiv:2307.15043). The spread across targets, from 86.6% down to 2.1%, is itself a sign that something fragile and surface-level was being exploited.
It also explains the cheapest attack anyone has published. Fine-tuning GPT-3.5 Turbo on ten examples, for less than $0.20 through the hosted API, produced a model that followed almost any harmful instruction (Qi et al., 2024, Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!, ICLR, arXiv:2310.03693). A few gradient steps cannot rewrite a deeply distributed policy. They can overwrite a shallow one.
The constructive half of the Qi et al. argument is a data augmentation: train on examples whose assistant turn begins harmful and then transitions into a refusal, so gradient pressure lands past the first few tokens and the model learns a route back from states it would otherwise never visit. Paired with a token-wise constrained fine-tuning objective that holds the early-token distribution in place, this improves robustness to prefilling and suffix attacks. It raises the number of tokens an attacker must control. It does not bound that number.
The new geometry is a budget curve
If there is no radius, what replaces \(\varepsilon\) as the knob the attacker turns? The answer measured across 2024 and 2025 is the sampling budget.
Best-of-N jailbreaking is the purest demonstration. Take one harmful request, apply random augmentations, capitalisation and character shuffling for text, pitch and speed shifts and background noise for audio, and resample until something lands. No gradients, no model internals, no cleverness. With 10,000 samples it reached a 78% attack success rate on Claude 3.5 Sonnet, and at least 52% on every language model tested (Hughes et al., 2024, Best-of-N Jailbreaking, NeurIPS 2025, arXiv:2412.03556).
If attempts were independent, cumulative success would be \(1 - (1-p)^N\) and would hit near-certainty for any \(p > 0\) once \(N \gg 1/p\). Augmented variants of one prompt are correlated, so the real curve is slower. What the measurements fit is a power law in the negative log of the failure rate:
with \(a\) and \(b\) estimated per model and modality. The paper's own validation is the detail that matters operationally: fitting on the first 1,000 samples and forecasting the rate at 10,000 gave a mean error of 4.6 percentage points across models and modalities. Robustness becomes a projection with error bars rather than a scalar.
Many-shot jailbreaking spends the budget inside one prompt instead of across many requests, stacking up to 256 faux dialogues in which the assistant complies. Its success also follows a power law, and that power law matches the scaling of benign in-context learning on the same models (Anthropic, 2024, Many-shot Jailbreaking). The attack rides the capability, which is why one prompt-based mitigation cut success from 61% to 2% while fine-tuning mitigations only delayed the attack.
Multimodality hands back the gradient
Discreteness was text's only structural protection. There is no small perturbation of a token, so attacks must search a combinatorial space, and GCG pays for that with hundreds of thousands of forward passes.
Vision and audio remove the constraint. Carlini and colleagues built adversarial images that drive aligned vision-language models into output they would never produce from text, and used the result to argue that existing text attacks are weak rather than that models are robust (Carlini et al., 2023, Are Aligned Neural Networks Adversarially Aligned?, NeurIPS, arXiv:2306.15447). Three properties stack for the attacker: the input is continuous so PGD applies at one backward pass per step; the dimensionality is enormous, roughly 339,000 free parameters in a 336 by 336 RGB patch against a 20-token suffix; and the perturbation is invisible, which defeats perplexity filtering and every other inspection-based defence that catches gibberish.
The same logic applies without gradients at all, which is why best-of-N's power laws hold across text, vision and audio alike.
[IMAGE: Horizontal log-scale cost axis, cheapest on the left. Markers placed along it: "benign-dataset fine-tune, accidental", "10 examples, $0.20, unalignment", "prefill assistant turn, free", "best-of-N, 10k samples, tens of dollars", "adversarial image, one PGD run", "GCG suffix, 256k forward passes", "universal jailbreak against a classifier cascade, thousands of human hours". Caption: "The attack surface sorted by price. Defences move items rightwards; none removes them."]
Where the defence moved to
Two families survived the move away from input-space thinking, and they moved in opposite directions.
Outwards, into layers. Constitutional classifiers put a separately trained model in front of and behind the generator. The classifiers are trained not on collected attacks but on synthetic data generated from a natural-language constitution describing permitted and restricted content, including data transformed to resemble known jailbreak styles. The streaming output classifier can halt a generation partway through, so an attacker must defeat the input screen and sustain the evasion across the whole response. Automated red-team success against the guarded production model fell from 86% to 4.4%, at a cost of 0.38 percentage points of extra refusals on real traffic and 23.7% additional inference compute (Sharma et al., 2025, Constitutional Classifiers, arXiv:2501.18837).
Inwards, into representations. Circuit breakers skip attack enumeration and train the model so that internal states on a trajectory towards harmful output are rerouted to states that cannot continue it, with a retain loss holding benign states fixed (Zou et al., 2024, Improving Alignment and Robustness with Circuit Breakers, NeurIPS, arXiv:2406.04313). Targeted latent adversarial training attacks the same surface during training, perturbing hidden activations adversarially and training the model to behave correctly under them; it beat the R2D2 adversarial-training baseline with orders of magnitude less compute while holding MMLU and MT-Bench flat (Sheshadri et al., 2024, arXiv:2407.15549).
The appeal of the inward move is coverage: one internal signature stands in for every input that produces it, including attacks nobody has invented. The limitation is that the mapping from a latent ball to a set of reachable inputs is unknown, which is exactly where the first adaptive-attack lesson landed again. Circuit breakers were reported robust to embedding-space attacks; Schwinn and Geisler removed constraints the original evaluation had imposed on that attack and reached a 100% attack success rate, raising ASR by more than 80 points over the published number (Schwinn & Geisler, 2024, Revisiting the Robust Alignment of Circuit Breakers, arXiv:2407.15902).
Seeing It in Motion
A cascaded defence and a budget-spending attacker are both loops, and the interesting behaviour is in how the loops interleave.
sequenceDiagram
participant A as Attacker
participant C1 as Cheap screen
participant C2 as Expensive classifier
participant M as Model
participant O as Output classifier
A->>C1: Augmented variant i
C1-->>A: Blocked, no signal returned
Note over A,C1: Attacker increments i and repeats
A->>C1: Variant i plus 400
C1->>C2: Escalated as suspicious
C2->>M: Allowed with elevated scrutiny
M->>O: Streaming tokens
O-->>A: Generation halted mid-response
Note over A,O: Partial output leaks a weak success signal
A->>C1: Variant tuned to the halt pointTwo things in that trace decide the outcome. The cheap screen returns no information when it blocks, so the attacker learns nothing from most of the budget. The output classifier, by halting mid-response, necessarily leaks a little: the attacker sees how far they got. Every defence that acts after generation starts faces that tradeoff, and the cascade's job is to make the uninformative path cover almost all traffic.
stateDiagram-v2
[*] --> Screened
Screened --> Refused: input classifier fires
Screened --> Generating: passes
Generating --> Halted: output classifier fires mid-stream
Generating --> Recovered: deepened alignment transitions to refusal
Generating --> Completed: no trigger
Halted --> [*]
Refused --> [*]
Recovered --> [*]
Completed --> [*]Recovered is the state that shallow alignment does not have. A model aligned only over its first tokens, once past them, has no path from Generating back to a refusal; the depth-augmentation training exists to create that edge.
Watch It Run
By the Numbers
| Quantity | Value | Source |
|---|---|---|
| CIFAR-10 robust accuracy, \(\ell_\infty\), \(\varepsilon = 8/255\) | 45.8% in 2018 (87.3% clean) | Madry et al., 2018 |
| Same, with generated-data augmentation | 70.69% (93.25% clean) | Wang et al., 2023, ICML |
| ICLR 2018 non-certified white-box defences relying on obfuscated gradients | 7 of 9; 6 broken fully, 1 partially | Athalye et al., 2018 |
| Defences broken by adaptive attacks, 2018 to 2020 | 13 | Tramèr et al., 2020 |
| LLM jailbreak and prompt-injection defences broken, 2025 | 12, most originally near-zero ASR, pushed above 90% | Nasr et al., 2025 |
| GCG suffix transfer ASR | 86.6% GPT-3.5, 66.0% PaLM-2, 47.9% Claude-1, 46.9% GPT-4, 2.1% Claude-2 | Zou et al., 2023 |
| Best-of-N text ASR at \(N = 10{,}000\) | 78% Claude 3.5 Sonnet; at least 52% on all LLMs tested | Hughes et al., 2024 |
| Best-of-N power-law forecast error, 1k to 10k samples | 4.6 percentage points, mean across models and modalities | Hughes et al., 2024 |
| Many-shot jailbreaking, prompt-based mitigation | 61% to 2% ASR; fine-tuning mitigations only delayed it | Anthropic, 2024 |
| Cost to unalign GPT-3.5 Turbo by fine-tuning | 10 examples, under $0.20 | Qi et al., 2024 |
| Constitutional classifiers, automated red team | 86% to 4.4% ASR; +0.38pp refusals; +23.7% compute | Sharma et al., 2025 |
| Red-team hours to first universal jailbreak on that system | None in 3,000+ hours by 183 people; one in about 3,700 more by 339 | Anthropic, 2025 |
| Constitutional Classifiers++ | about 40x cheaper classifier stage, roughly 1% total overhead, 0.05% refusals on harmless traffic | Cunningham, Wei et al., 2026 |
| Same, red-team result | 1,700+ hours over 198,000 attempts; one high-risk finding; 0.005 detections per thousand queries | Anthropic, 2026 |
| Circuit breakers under an unconstrained embedding attack | 100% ASR, more than 80 points above the published evaluation | Schwinn & Geisler, 2024 |
Sources: every figure above is from the linked primary source. The CIFAR-10 numbers are AutoAttack evaluations at \(\varepsilon = 8/255\). Refusal and compute overheads for the classifier systems are vendor-reported measurements on that vendor's own production traffic and have not been independently reproduced; treat them as the best available figures rather than as audited ones.
A Concrete Example
Here is the arithmetic a defender should actually run. All token prices below are illustrative, chosen to be in the right range for a mid-tier frontier model in 2026; substitute your own.
Step 1: measure two points, not one. You have a guarded endpoint and a behaviour set of 100 harmful requests. Send each request once, unaugmented. Four succeed.
Now send 100 augmented variants per request. Thirty-one of the hundred behaviours succeed at least once.
Step 2: forecast the budget you did not spend. At \(N = 10{,}000\):
A 4% single-shot success rate implies roughly 97% at ten thousand samples. The headline "96% of our refusals held" and "the attacker wins 97% of the time" are the same measurement at different budgets.
Step 3: price it from the attacker's side. Suppose each attempt is 400 input and 300 output tokens, at $3 per million input and $15 per million output:
Ten thousand samples against one target behaviour is $57 and, at three requests per second, under an hour of wall clock. That is the real number. Not "the model is 96% robust", but "one behaviour, 97% likely, $57, 55 minutes".
Step 4: work out what a defence has to deliver. To hold ASR at 10,000 samples down to 5%, you need \(aN^b = -\log(0.95) = 0.0513\), so \(a = 0.0513 / 82.7 = 6.2 \times 10^{-4}\). Since \(a \approx \text{ASR}(1)\) for small rates, single-shot success must fall from 4% to about 0.062%.
That is a 65-fold reduction in per-sample success, and it buys a move from 96.6% to 5% at that one budget. Cut the attacker's budget requirement by 65x and they respond by spending 65x more, which at $0.0057 per sample is $3,700 and a weekend.
Step 5: check the control you were counting on. Rate limiting looks like the answer, so test it. Cap one account at 200 requests per day against this endpoint:
Forty percent per account per day. The cap did not help much, because the curve is sublinear in exactly the regime the cap puts you in. What does help is detecting that 200 near-duplicate requests came from one source, which is a correlation control rather than a volume control. Walking the arithmetic is what tells you that; intuition says the opposite.
[IMAGE: Line chart, log-x. X axis: sampling budget N from 1 to 100,000. Y axis: attack success rate 0 to 100%. Curve A fitted with a = 0.0408, b = 0.479, annotated at N=1 (4%), N=200 (40%), N=10,000 (96.6%). Curve B is the same exponent with a reduced 65-fold, annotated at N=10,000 (5%) and N=650,000 (still rising). Caption: "A 65x reduction in single-shot success moves the curve right, not down."]
Where It Breaks
Defence evaluations are lower bounds on attack success, always
The single most reliable finding in this literature is that a defence's own evaluation understates attack success. It happened in 2018, in 2020, to circuit breakers in 2024, and to twelve defences in 2025. The mechanism is not dishonesty. It is that the authors optimised the defence and not the attack, and that the attack an author reaches for is the attack that exists, not the one designed against their specific mechanism.
The practical consequence: when reading a defence paper, treat its ASR as the floor. When writing one, report what an attacker who read your method would do, with the compute you gave them.
"Flake" has an equivalent here, and it is called "the attack was weak"
Jain and colleagues argued that discrete token optimisation is weak and expensive enough that simple defences such as perplexity filtering work better for LLMs than their analogues did in vision (Jain et al., 2023, Baseline Defenses for Adversarial Attacks Against Aligned Language Models, arXiv:2309.00614). That read the 2023 evidence correctly and aged badly anyway, because it described one threat model. A perplexity filter assumes attacks look like gibberish. Best-of-N variants are fluent, and images carry no perplexity at all.
[IMAGE: Timeline of five defence cohorts (ICLR 2018, ICLR/ICML/NeurIPS 2018-2020, circuit breakers 2024, twelve LLM defences 2025), each drawn as a paired bar: published attack success rate in teal next to post-adaptive-attack rate in rose. Every pair shows the rose bar far taller. Caption: "No defence paper's published ASR has yet turned out to be pessimistic."]
Graders decide what counts as a success
ASR is whatever the judge scores as harmful. Jailbreak papers have systematically overstated effectiveness because weak graders count empty or useless compliance, which is the inflation StrongREJECT was built to correct with a grader fine-tuned against human judgement (Souly et al., 2024, A StrongREJECT for Empty Jailbreaks, NeurIPS Datasets and Benchmarks, arXiv:2402.10260). Fit a power law to a bad grader and you extrapolate a bad measurement with impressive precision.
Layers correlate, so depth multiplies less than it looks
Defence in depth assumes independent failures. When an input classifier, an output classifier and the model's own refusal training all derive from related rules and related synthetic data, an attack that defeats one is likelier than chance to defeat the others. The classifiers are also models, so prompt-level evasion and distribution shift apply to them directly, and their training pipelines are themselves a target: Anthropic's own safeguards research has demonstrated backdoors planted through a classifier's fine-tuning data (Anthropic Alignment Science, 2026, Poisoning Fine-tuning Datasets of Constitutional Classifiers).
Overrefusal aggregates hide the people who notice
A 0.05% refusal rate on harmless production traffic is an excellent number and it is an average. The rate for a toxicologist, a security engineer or an infectious-disease researcher is not 0.05%, because their legitimate work sits against the policy boundary. Aggregate overrefusal metrics are dominated by ordinary traffic and are therefore least informative about the users most affected.
Open weights void most of the warranty
Refusing prefill, hiding log-probabilities, clamping temperature, rate limiting and detecting correlated request streams are all properties of a serving stack rather than of a model. Download the weights and the attacker has unconstrained gradient descent, arbitrary prefill, full sampling control and the ability to fine-tune the alignment off for the price of a GPU hour. A robustness claim that does not state which interface it assumes is not a comparable measurement.
Alternative Designs
| Design | How it works | Key advantage | Key limitation | Best when |
|---|---|---|---|---|
| Refusal training (safety SFT and RLHF) | Preference data teaches the model to decline | Free at inference; no added latency | Shallow in token depth; removable by fine-tuning for cents | Always, as the base layer |
| Deepened alignment | Recovery examples plus token-wise constrained fine-tuning move gradient past the first tokens | Closes prefill, suffix and decoding attacks at once | Raises the token count an attacker must control without bounding it | You control the model and its fine-tuning |
| Input and output classifier layers | Separate models trained from a natural-language constitution screen traffic | Updatable in hours; catches attacks in any modality | Compute and refusal tax; the classifier is itself attackable | Hosted deployment with policy-defined harms |
| Cascaded classifiers | A cheap screen escalates only the suspicious tail | Roughly 40x less classifier compute than a single expensive stage | Strength still set by the expensive stage; escalation logic is a new surface | Production traffic where benign dominates |
| Representation-level defences | Reroute or adversarially train internal states associated with harmful behaviour | Covers attack forms nobody has invented | Coverage defined by example sets; latent balls need not map to reachable inputs | Open-weight releases, where input controls are unavailable |
| Deliberative alignment | Train the model to reason over its safety specification before answering | Improves jailbreak robustness and overrefusal together | Costs reasoning tokens; the specification becomes the attack target | Reasoning models with a token budget to spend |
| Interface restriction | No prefill, hidden log-probabilities, clamped sampling, correlation detection | Directly attacks the budget curve rather than \(p\) | Unavailable for open weights; nothing touches offline transfer attacks | Any hosted API |
| Certified defences | Prove no input in a formal set changes the output | A real guarantee | No formal input set exists for natural language | Vision and other metric input spaces |
[IMAGE: Two-axis scatter of the eight defence families. X axis: effect on single-shot success probability p, from none to large. Y axis: effect on the cost of a sample, from none to large. Interface restriction sits alone high on the Y axis; classifier layers and representation defences cluster on the X axis; certified defences sit off-chart in a greyed box labelled "no formal input set for text". Caption: "Most defences buy a lower p. Only one family makes each attempt more expensive."]
The row that does the most work in practice is the last-but-one. Everything else reduces \(p\); interface restriction changes what \(N\) costs. Given a sublinear budget curve, moving \(N\) is often the better investment, and it is the one most robustness papers do not measure.
How It Is Used in Practice
The operating pattern that has emerged at the frontier labs is a stack, not a defence: safety training in the weights, a constitution-derived classifier layer around them, interface restrictions on the serving path, and a standing red-team and bug-bounty programme whose output is the evidence that the stack works.
The economics of that stack improved sharply between 2025 and 2026, and the improvement was in the cost column rather than the robustness column. The first production constitutional classifier system cost 23.7% extra compute. The second, built around a two-stage cascade with lightweight screens escalating to expensive classifiers, plus efficient linear probes ensembled alongside, cut the classifier stage by roughly 40x to about 1% overhead and brought harmless-traffic refusals to 0.05% (Cunningham, Wei et al., 2026, Constitutional Classifiers++, arXiv:2601.04603). Over 1,700 cumulative red-team hours across 198,000 attempts, one high-risk vulnerability was found, a detection rate of 0.005 per thousand queries (Anthropic, 2026, Next-generation Constitutional Classifiers).
Read that as an engineering result about cost, which is what it is. A defence at 23.7% overhead is a tax on every token the business sells; at 1% it is a rounding error, which means it can be deployed everywhere rather than on the highest-risk surfaces only. Coverage bought by cost reduction is a larger safety gain than most percentage-point improvements in ASR.
[IMAGE: Stacked horizontal bar comparing two generations of a safeguard system on three axes, each normalised: inference compute overhead (23.7% versus about 1%), harmless-traffic refusal rate (0.38 percentage points versus 0.05%), and red-team hours survived per universal jailbreak found. Caption: "Between 2025 and 2026 the safeguard improved mostly in price, and price is what determines coverage."]
The reporting discipline matters as much as the stack. A claim like "no red-teamer found a universal jailbreak" is evidence about a specific search, by specific people, with specific incentives, over a specific number of hours. The Constitutional Classifiers sequence demonstrated the limit of that claim on its own system: more hours from a larger crowd produced the universal jailbreak the earlier programme had not found, against an unchanged defence. Publishing both results is what makes the first one worth anything.
Insights Worth Remembering
-
An attack success rate without a budget is not a measurement. The same system reports 4% at one sample and 97% at ten thousand. Every ASR in a paper, a dashboard or a vendor claim should be read as a function evaluated at an unstated point, and the first question is always what the point was.
-
Reducing per-sample success is multiplicative, not protective. With a power-law exponent near 0.5, a 65-fold cut in single-shot success moves the curve sideways. It spends the attacker's budget rather than causing their failure.
-
Four famous attacks are one vulnerability. Prefilling, temperature, adversarial suffixes and fine-tuning all exploit the fact that alignment lives in the first few output tokens. That is why ten fine-tuning examples suffice and why GCG optimises for "Sure, here is" instead of the harmful content.
-
Adding a modality is adding an attack surface with better economics. An image gives the attacker continuity, roughly 339,000 free parameters and imperceptibility in one step. Multimodal capability ships ahead of multimodal safety coverage every time, and the gap is widest right after launch.
-
A defence's own evaluation is a floor. 2018, 2020, circuit breakers in 2024, twelve defences in 2025. There is no recorded case of a defence paper's published ASR turning out to be pessimistic. Plan your reading, and your own evaluations, around that asymmetry.
-
Defence in depth is weaker than the product of its layers. Correlated training data and shared rules mean correlated failures. If your input classifier, output classifier and refusal training all descend from the same constitution, do not multiply their miss rates together and call it a bound.
-
The cheapest robustness gain available is usually a cost reduction, not an ASR reduction. Dropping safeguard overhead from 23.7% to 1% means the safeguard runs everywhere, which changes realised risk more than a few ASR points on surfaces it already covered.
-
Robustness claims expire. Attack cost falls with every published technique, cheaper model and open-weight surrogate. "Robust as of early 2026" is the only honest tense, and a dated claim with a budget attached is worth more than an undated one with a bigger number.
Open Questions
Can any useful formal threat model be defined for natural language? What is known: no metric, projection or closed form currently exists for the set of prompts eliciting a capability, and all certification machinery needs one. What is speculation: whether a coarse but formal surrogate, say certified robustness over a defined set of semantic transformations, would be tight enough to be worth proving. Nobody has shown that such a set covers the attacks that matter.
How correlated are the layers in a real stack? Measured: individual layer miss rates. Not measured, publicly: the joint distribution. A single experiment reporting conditional miss rates, the probability the output classifier misses given the input classifier missed, would tell operators more than another round of per-layer numbers, and the data to run it exists inside the labs.
Does representation-level defence generalise to attacks outside its behaviour set? The attack-agnostic claim is about attack form, and the evidence supports it there. Whether it extends to harmful behaviours absent from the training set is untested, and the Schwinn and Geisler result shows the evaluation bar is high enough that early positive results should be treated as provisional.
Do reasoning models change the shape of the curve or only its constants? Deliberative alignment improves jailbreak robustness and overrefusal together (Guan et al., 2024, Deliberative Alignment, arXiv:2412.16339), which is a genuine Pareto move. Whether spending more thinking tokens flattens the power-law exponent \(b\), rather than just lowering \(a\), is an open empirical question, and it is the one that decides whether test-time compute is a defence or a discount.
What is the right unit for publishing an attacker-cost claim? Samples, hours, dollars and FLOPs all appear in the literature and none converts cleanly to the others. Until the field settles on a unit, cost-based robustness claims are not comparable across papers, which is the same problem ASR-without-a-budget has, one level up.
Sources and Further Reading
- Szegedy, C., Zaremba, W., Sutskever, I., Bruna, J., Erhan, D., Goodfellow, I., & Fergus, R. (2014). "Intriguing Properties of Neural Networks." ICLR. arXiv:1312.6199
- Madry, A., Makelov, A., Schmidt, L., Tsipras, D., & Vladu, A. (2018). "Towards Deep Learning Models Resistant to Adversarial Attacks." ICLR. arXiv:1706.06083
- Athalye, A., Carlini, N., & Wagner, D. (2018). "Obfuscated Gradients Give a False Sense of Security: Circumventing Defenses to Adversarial Examples." ICML. arXiv:1802.00420
- Tramèr, F., Carlini, N., Brendel, W., & Madry, A. (2020). "On Adaptive Attacks to Adversarial Example Defenses." NeurIPS. arXiv:2002.08347
- Wang, Z., Pang, T., Du, C., Lin, M., Liu, W., & Yan, S. (2023). "Better Diffusion Models Further Improve Adversarial Training." ICML. arXiv:2302.04638
- Carlini, N., Nasr, M., Choquette-Choo, C. A., Jagielski, M., Gao, I., et al. (2023). "Are Aligned Neural Networks Adversarially Aligned?" NeurIPS. arXiv:2306.15447
- Zou, A., Wang, Z., Carlini, N., Nasr, M., Kolter, J. Z., & Fredrikson, M. (2023). "Universal and Transferable Adversarial Attacks on Aligned Language Models." arXiv:2307.15043
- Jain, N., Schwarzschild, A., Wen, Y., Somepalli, G., Kirchenbauer, J., et al. (2023). "Baseline Defenses for Adversarial Attacks Against Aligned Language Models." arXiv:2309.00614
- Qi, X., Zeng, Y., Xie, T., Chen, P., Jia, R., Mittal, P., & Henderson, P. (2024). "Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!" ICLR. arXiv:2310.03693
- Qi, X., Panda, A., Lyu, K., Ma, X., Roy, S., Beirami, A., Mittal, P., & Henderson, P. (2024). "Safety Alignment Should Be Made More Than Just a Few Tokens Deep." arXiv:2406.05946
- Zou, A., Phan, L., Wang, J., Duenas, D., Lin, M., Andriushchenko, M., et al. (2024). "Improving Alignment and Robustness with Circuit Breakers." NeurIPS. arXiv:2406.04313
- Schwinn, L., & Geisler, S. (2024). "Revisiting the Robust Alignment of Circuit Breakers." arXiv:2407.15902
- Sheshadri, A., Ewart, A., Guo, P., Lynch, A., Wu, C., et al. (2024). "Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs." arXiv:2407.15549
- Souly, A., Lu, Q., Bowen, D., Trinh, T., Hsieh, E., et al. (2024). "A StrongREJECT for Empty Jailbreaks." NeurIPS Datasets and Benchmarks. arXiv:2402.10260
- Hughes, J., Price, S., Lynch, A., Schaeffer, R., Barez, F., et al. (2024). "Best-of-N Jailbreaking." NeurIPS 2025. arXiv:2412.03556
- Guan, M. Y., Joglekar, M., Wallace, E., Jain, S., Barak, B., et al. (2024). "Deliberative Alignment: Reasoning Enables Safer Language Models." arXiv:2412.16339
- Sharma, M., Tong, M., Mu, J., Wei, J., Kruthoff, J., et al. (2025). "Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming." arXiv:2501.18837
- Nasr, M., Carlini, N., Sitawarin, C., Schulhoff, S., et al. (2025). "The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against LLM Jailbreaks and Prompt Injections." arXiv:2510.09023
- Cunningham, H., Wei, J., et al. (2026). "Constitutional Classifiers++: Efficient Production-Grade Defenses against Universal Jailbreaks." arXiv:2601.04603
- Anthropic (2024). "Many-shot Jailbreaking." anthropic.com/research/many-shot-jailbreaking
- Anthropic (2025). "Constitutional Classifiers: Defending against Universal Jailbreaks." anthropic.com/research/constitutional-classifiers
- Anthropic (2026). "Next-generation Constitutional Classifiers." anthropic.com/research/next-generation-constitutional-classifiers
- Anthropic Alignment Science (2026). "Poisoning Fine-tuning Datasets of Constitutional Classifiers." alignment.anthropic.com/2026/backdooring-classifiers
Free to read, no ads, no sign-up. If it was useful you can buy me a coffee.