Adversarial Robustness advanced 9 min read 6 flashcards

Embedding-Space Attacks and Representation-Level Defences

How relaxing the discrete token constraint restores the full strength of gradient attacks, why circuit breakers and latent adversarial training operate on representations instead of inputs, and the evaluation failure that gave one of them a 100% attack success rate.

Text is discrete, and that discreteness is the only reason language models are not as trivially attackable as image classifiers. There is no small perturbation of a token; the nearest neighbour of a word is another word. Attacks like GCG pay for this with hundreds of thousands of forward passes to search a combinatorial space, as Black-Box and Transfer Attacks describes.

Remove the constraint and the advantage evaporates. If an attacker can supply embeddings rather than token IDs, which anyone holding open weights can, the input space is \(\mathbb{R}^{n \times d}\) and projected gradient descent applies exactly as it does to pixels.

The continuous relaxation

A token sequence enters the model as \(E = [\,e(t_1), \ldots, e(t_n)\,]\) with \(e(\cdot)\) the embedding lookup. An embedding-space attack optimises \(E\) directly:

\[\min_{\delta} \; -\log p_\theta\big(y^\star \mid E + \delta\big), \qquad \|\delta\| \le \varepsilon \;\; \text{(or unconstrained)},\]

solved with the same PGD loop as vision. Each step costs one forward and one backward pass, against GCG's hundreds of candidate evaluations per step, so the attack is orders of magnitude cheaper than discrete search and strictly stronger, because the token vocabulary is a finite subset of the space being searched.

The resulting \(\delta\) is usually not a real prompt, which is why this is a defence evaluation tool rather than a deployment threat for a hosted model. It is the right threat model for open weights, and the right stress test for any defence that claims to be attack-agnostic.

Defending at the representation, not the input

If attacks can enter through the representation, defences can live there too. Circuit breakers take this literally: instead of enumerating attacks, train the model so that internal states on a trajectory towards harmful output are rerouted to states that cannot continue it. Representation Rerouting adds a loss that pushes hidden states in a harmful-behaviour set away from their original direction while a retain loss holds benign states fixed (Zou et al., 2024, Improving Alignment and Robustness with Circuit Breakers, NeurIPS, arXiv:2406.04313). The claim is attack-agnostic robustness, because the intervention targets the behaviour's internal signature rather than any input pattern.

Targeted latent adversarial training attacks the same surface during training: perturb hidden activations adversarially, then train the model to behave correctly under those perturbations. It beats the R2D2 adversarial-training baseline on jailbreak robustness with orders of magnitude less compute, and the authors report MMLU and MT-Bench held flat (Sheshadri et al., 2024, Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs, arXiv:2407.15549). The appeal over input-space adversarial training is coverage: a latent perturbation stands in for a whole family of inputs that would produce that activation, including inputs nobody has invented.

When it breaks

The attack that evaluates the defence must itself be adaptive. Circuit breakers were reported robust to embedding-space attacks. A follow-up showed that straightforward changes to the embedding attack, mainly removing the constraints the original evaluation imposed, reached a 100% attack success rate against circuit-breaker models, raising ASR by more than 80 points over the published evaluation (Schwinn & Geisler, 2024, Revisiting the Robust Alignment of Circuit Breakers, arXiv:2407.15902). This is the vision-era lesson of Robustness Evaluation That Means Something, repeated verbatim in a new input space.

Representation sets are defined by examples. Both methods need a harmful-behaviour set and a retain set. Behaviour outside the harmful set is not rerouted, and behaviour mistakenly inside the retain set is protected, so coverage is an annotation problem wearing a mechanistic costume.

Latent perturbations are a proxy, not a superset. A perturbation at layer \(\ell\) need not correspond to any reachable input, so training against it can buy robustness in directions no attacker can enter while leaving reachable directions untouched. The mapping from latent balls to input sets is unknown, and that is the honest limitation.

Rerouting shows up as degraded generations, not clean refusals. A model interrupted mid-representation often emits incoherent or truncated text rather than a policy-shaped refusal, which is harder to audit and harder to explain to a user who hit it on a false positive.

References and further reading

Every source this page cites, in the order it cites them. All of them open in a new tab.

  1. Zou et al., 2024, Improving Alignment and Robustness with Circuit Breakers, NeurIPS, arXiv:2406.04313 arxiv.org
  2. Sheshadri et al., 2024, Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs, arXiv:2407.15549 arxiv.org
  3. Schwinn & Geisler, 2024, Revisiting the Robust Alignment of Circuit Breakers, arXiv:2407.15902 arxiv.org
Check yourself

6 flashcards for this concept

Click a card to reveal the answer.

Drill the whole track