Adversarial Robustness intermediate 8 min read 6 flashcards

Attack Scaling and the Sampling Budget

Why attack success against a language model is a function of how many times the attacker is allowed to try, how many-shot and best-of-N jailbreaking turn that into a predictable power law, and what it does to any robustness number reported at a single budget.

Take one harmful request, apply random capitalisation and character shuffling to it, and send 10,000 variants. Nothing is optimised, no gradient is computed, and no model internals are touched. On Claude 3.5 Sonnet that procedure reached a 78% attack success rate, and across every language model tested it reached at least 52% (Hughes et al., 2024, Best-of-N Jailbreaking, NeurIPS 2025, arXiv:2412.03556). The interesting part is not the number. It is that the number was predictable before the experiment finished.

In the norm-ball world of Adversarial Examples and the Threat Model, robustness is a property of a model at a radius \(\varepsilon\): either a perturbation exists inside the ball or it does not. Against a language model there is no radius, and an attacker with a sampling budget does not need to find the single best attack. They need to find any attack, once.

Success is a budget curve, not a point

Write \(p\) for the probability that a single sampled attack succeeds and \(N\) for the attacker's budget. If attempts were independent, the cumulative success rate would be

\[\text{ASR}(N) = 1 - (1 - p)^N,\]

which reaches near-certainty for any \(p > 0\) once \(N \gg 1/p\). Real attack sampling is not independent, because augmentations of one prompt are correlated, so the curve is slower than that geometric bound. What the measurements show instead is that the negative log of the failure rate grows as a power of \(N\):

\[-\log\big(1 - \text{ASR}(N)\big) \approx a N^{b},\]

with \(a\) and \(b\) fit per model and modality. Hughes et al. verified this by fitting on the first 1,000 samples and forecasting the ASR at 10,000, with a mean error of 4.6 percentage points across models and modalities. A defender can therefore be told, from a short experiment, roughly what an attacker with a hundred times the budget will achieve.

The same shape appears when the budget is spent inside one prompt rather than across many. Stacking faux dialogues in which the assistant complies, up to 256 of them, drives harmful response rates up along a power law that matches the scaling of benign in-context learning on the same models (Anthropic, 2024, Many-shot Jailbreaking). That matching is the uncomfortable finding: the attack rides the model's in-context learning ability, so improving that ability improves the attack.

What this breaks about reported robustness

An attack success rate is meaningless without the budget that produced it. "4% ASR" from a defence paper and "78% ASR" from an attack paper are often the same system measured at \(N = 1\) and \(N = 10{,}000\).

Three consequences follow. Defences that reduce \(p\) by a constant factor buy a multiplicative increase in required budget, not immunity; to halve \(p\) when \(b \approx 0.5\) costs the attacker roughly four times the samples. Defences that are stateless across requests cannot see the budget being spent at all, which is why per-request filtering and cross-request rate limiting are different controls. And a benchmark that fixes \(N\) ranks defences only at that budget, so a defence tuned to it can look strong and fail at \(10N\).

When it breaks

Correlated samples flatten the curve. Augmentations drawn from one template explore a narrow region. An attacker who diversifies templates gets a steeper curve than the power-law fit from a single template predicts, so the forecast is a lower bound on a creative attacker, not a ceiling.

The grader is in the loop. ASR is whatever the judge model scores as harmful. Jailbreak papers have systematically overstated effectiveness because weak graders count empty or useless compliance as success, which is the gap StrongREJECT was built to close with a fine-tuned grader scored against human judgement (Souly et al., 2024, A StrongREJECT for Empty Jailbreaks, NeurIPS Datasets and Benchmarks, arXiv:2402.10260). A power law fit to a bad grader extrapolates a bad measurement precisely.

Budget is not the only axis. Many-shot attacks need context length, best-of-N needs throughput, and discrete optimisation needs forward passes on a surrogate. Three attackers with the same dollar budget have different reachable ASRs depending on which resource the target is cheap in.

The defender's budget scales too. Two-stage classifier cascades exploit the same asymmetry in reverse, screening all traffic cheaply and escalating only suspicious exchanges, which is how the compute cost of a safeguard layer fell by roughly 40x between 2025 and 2026 (Cunningham, Wei et al., 2026, Constitutional Classifiers++, arXiv:2601.04603). Robustness under budget is a race between two cost curves, covered in Attacker Cost as the Robustness Metric.

References and further reading

Every source this page cites, in the order it cites them. All of them open in a new tab.

  1. Hughes et al., 2024, Best-of-N Jailbreaking, NeurIPS 2025, arXiv:2412.03556 arxiv.org
  2. Anthropic, 2024, Many-shot Jailbreaking anthropic.com
  3. Souly et al., 2024, A StrongREJECT for Empty Jailbreaks, NeurIPS Datasets and Benchmarks, arXiv:2402.10260 arxiv.org
  4. Cunningham, Wei et al., 2026, Constitutional Classifiers++, arXiv:2601.04603 arxiv.org
Check yourself

6 flashcards for this concept

Click a card to reveal the answer.

Drill the whole track