Adversarial Robustness advanced 7 min read 6 flashcards

Cross-Modal Attack Surfaces

Why adding vision or audio to a language model hands back the continuous input space and the gradient attack that came with it, and why the resulting attacks are both stronger and cheaper than their text equivalents.

Adversarial examples were an image problem for a decade before they were a language problem (Szegedy et al., 2014, Intriguing Properties of Neural Networks, ICLR, arXiv:1312.6199). Multimodal models return language models to that setting. Pixels are continuous, differentiable and perceptually redundant, so an attacker gets a genuine \(\varepsilon\)-ball to search and a human-imperceptible budget to hide in.

Carlini et al. made the point by constructing adversarial images that drive aligned vision-language models into harmful output they would never produce from text alone, and used the result to argue that existing text attacks are weak rather than that the models are robust (Carlini et al., 2023, Are Aligned Neural Networks Adversarially Aligned?, NeurIPS, arXiv:2306.15447). The conclusion is structural: alignment trained in one modality does not transfer to a jointly embedded one.

Why the image channel is the cheap one

Three properties stack in the attacker's favour.

The input is continuous, so PGD applies directly and each step costs one backward pass, with no combinatorial search over a vocabulary.

The dimensionality is enormous. A \(336 \times 336\) RGB patch carries roughly 339,000 free parameters against the 20 tokens a GCG suffix optimises, and attack strength grows with the dimension of the space being searched.

The perturbation is invisible, which removes the only cheap defence text has. A gibberish suffix can be caught by a perplexity filter; an \(\ell_\infty\) perturbation of \(4/255\) cannot be caught by looking at the image.

Audio behaves the same way, with the additional property that augmentations which are inaudible to a listener, speed and pitch shifts, background noise, change the token sequence a speech encoder produces. Best-of-N jailbreaking exploits exactly this, and the power-law scaling it reports holds across text, vision and audio inputs (Hughes et al., 2024, arXiv:2412.03556).

When it breaks

Defending the encoder is not defending the system. Hardening an image encoder against classification attacks says nothing about perturbations optimised against a downstream language objective, because the attacker's loss is the generative one, not the encoder's.

Input preprocessing is a weak and well-studied defence. JPEG compression, random resizing and additive noise each raise attack cost and each fell to adaptive attacks in the vision literature, most of them because they mask gradients rather than remove the vulnerability (Athalye et al., 2018, Obfuscated Gradients Give a False Sense of Security, ICML, arXiv:1802.00420).

Output-side safeguards see the harm, not the vector. A classifier reading the generated response can catch harmful output regardless of which modality carried the attack, which is the main reason layered defences are attractive here. It also means the only signal the defender gets is after the model has already been steered.

Modality coverage lags capability. Safety training data, red-team coverage and classifier training sets are overwhelmingly textual. Every new input channel, documents as images, video, screen captures for computer-use agents, opens a surface where the attack cost is lower and the evidence base is thinner.

References and further reading

Every source this page cites, in the order it cites them. All of them open in a new tab.

  1. Szegedy et al., 2014, Intriguing Properties of Neural Networks, ICLR, arXiv:1312.6199 arxiv.org
  2. Carlini et al., 2023, Are Aligned Neural Networks Adversarially Aligned?, NeurIPS, arXiv:2306.15447 arxiv.org
  3. Hughes et al., 2024, arXiv:2412.03556 arxiv.org
  4. Athalye et al., 2018, Obfuscated Gradients Give a False Sense of Security, ICML, arXiv:1802.00420 arxiv.org
Check yourself

6 flashcards for this concept

Click a card to reveal the answer.

Drill the whole track