Adversarial Robustness advanced 8 min read 6 flashcards

Shallow Safety Alignment and Generation-Interface Attacks

Why refusal behaviour concentrated in a model's first few output tokens explains prefilling, decoding-parameter and fine-tuning attacks at once, and what deepening alignment past those tokens actually changes.

Ten training examples and less than $0.20 of fine-tuning API spend were enough to strip the safety behaviour out of GPT-3.5 Turbo and leave a model that would follow almost any harmful instruction (Qi et al., 2024, Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!, ICLR, arXiv:2310.03693). The same paper found that fine-tuning on entirely benign, commonly used datasets degraded safety too, just less.

That result looks like a different problem from an adversarial suffix, which looks like a different problem from forcing the assistant turn to begin with "Sure, here is". The argument of Safety Alignment Should Be Made More Than Just a Few Tokens Deep is that all three are the same problem seen from different interfaces (Qi et al., 2024, arXiv:2406.05946).

Alignment as a prefix policy

Refusal in an aligned chat model is implemented mostly as a shift in the output distribution over the first handful of tokens. Condition on "I cannot" and the continuation stays refusing; condition on "Sure, here is a" and the model's conditional distribution for the rest of the response is close to the unaligned base model's, because almost no alignment gradient ever reached that region of the sequence.

Formally, alignment moves \(p_\theta(y_{1:k} \mid x)\) for small \(k\) while leaving \(p_\theta(y_{k+1:T} \mid x, y_{1:k})\) nearly unchanged. Every interface that lets an attacker choose \(y_{1:k}\) therefore bypasses alignment without touching weights:

  • Prefilling. APIs and open-weight inference stacks that accept a partial assistant turn hand the attacker \(y_{1:k}\) directly.
  • Decoding parameters. Raising temperature, or sampling far down the tail with high top-\(k\), makes the model draw a non-refusing first token by chance; the budget argument in Attack Scaling and the Sampling Budget then does the rest.
  • Adversarial suffixes. GCG-style optimisation maximises \(\log p_\theta(y^\star_{1:k} \mid x \oplus s)\) for an affirmative target prefix \(y^\star\). The objective is explicitly a prefix objective, which is why the attack is cheap relative to optimising a whole harmful response.
  • Fine-tuning. A few gradient steps on compliant openings overwrite exactly the shallow region where alignment lives, which is why ten examples suffice.

Deepening the alignment

The constructive half of the argument is a data augmentation: train on examples whose assistant turn starts harmful and then recovers into a refusal, so that gradient pressure lands at positions \(k+1\) and beyond and the model learns a transition back to refusal from mid-response states it would otherwise never see. Qi et al. report that this improves robustness to prefilling and suffix attacks, and that a token-wise constrained fine-tuning objective, which holds the early-token distribution in place while allowing later tokens to move, makes the model meaningfully harder to unalign by fine-tuning.

Deliberative alignment pushes in the same direction from another angle, training the model to reason over its safety specification before answering, so that the decision is made in a long chain rather than in token one (Guan et al., 2024, Deliberative Alignment, arXiv:2412.16339). See /learn/deliberative-alignment.

When it breaks

Depth is not a certificate. Deepened alignment raises the number of tokens an attacker must control. An attacker who can prefill 50 tokens instead of five is back where they started, and nothing in the method bounds that.

Fine-tuning access defeats everything downstream of weights. Constrained objectives apply only if the defender controls the fine-tuning run. An open-weight release hands the attacker unconstrained gradient descent, and no amount of pre-release alignment survives it.

The open-weight and API threat models diverge here. Hiding log-probabilities, refusing prefill, and clamping temperature are real controls for a hosted model and not available at all for a downloaded one. Robustness claims that do not say which interface they assume are not comparable.

Recovery training has a utility cost. Teaching a model to break off mid-response and refuse makes it more likely to abandon legitimate responses that merely look risky, which shows up as overrefusal on benign-but-sensitive prompts rather than as a benchmark regression.

References and further reading

Every source this page cites, in the order it cites them. All of them open in a new tab.

  1. Qi et al., 2024, Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!, ICLR, arXiv:2310.03693 arxiv.org
  2. Qi et al., 2024, arXiv:2406.05946 arxiv.org
  3. Guan et al., 2024, Deliberative Alignment, arXiv:2412.16339 arxiv.org
Check yourself

6 flashcards for this concept

Click a card to reveal the answer.

Drill the whole track