Layered Defences and the Cost of Safeguards
What a classifier layer in front of and behind a model actually buys against universal jailbreaks, measured in success rate, false refusals and compute overhead, and why cascading cheap screens before expensive ones is the engineering that made it deployable.
An automated red-team suite got harmful content out of an undefended production model 86% of the time. Put constitutional classifiers in front of and behind the same model and that fell to 4.4%, at the price of a 0.38 percentage-point increase in refusals on real production traffic and 23.7% additional inference compute (Sharma et al., 2025, Constitutional Classifiers, arXiv:2501.18837; Anthropic, 2025, Constitutional Classifiers).
Those four numbers are the whole shape of the tradeoff: a large reduction in attack success, a small but real tax on legitimate users, a compute bill, and nothing resembling a guarantee.
What the layer is actually doing
A safeguard layer is a separate model trained from natural-language rules rather than from collected attack examples. A constitution describing what is and is not permitted is used to generate synthetic training data spanning both, including data deliberately transformed to resemble known jailbreak styles, and input and output classifiers are trained on it. Two properties follow from that design.
It is updatable on a different clock than the model. A new rule is a text edit plus a retraining run of a small classifier, measured in hours, against weeks for an alignment retrain. That decoupling is the main operational argument for the layer, and it is the same argument made in /learn/guardrail-classifiers-and-content-filtering.
It targets the universal jailbreak rather than the single response. A streaming output classifier can halt a generation partway through, so the attacker must defeat the input screen and sustain the evasion across the whole response. Raising the cost of extracting detailed harmful content across many target queries is the stated objective, not driving single-prompt ASR to zero.
The cost curve matters as much as the success rate
A defence that doubles inference cost does not ship. The second generation of the system attacked that number directly with a two-stage cascade: a cheap classifier screens all traffic and escalates only suspicious exchanges to an expensive one, with efficient linear probes ensembled alongside. The reported result is roughly a 40x reduction in classifier compute against the baseline exchange classifier, about 1% additional compute overall, and a 0.05% refusal rate on harmless production traffic, an 87% drop from the first system (Cunningham, Wei et al., 2026, Constitutional Classifiers++, arXiv:2601.04603; Anthropic, 2026, Next-generation Constitutional Classifiers).
The economics are the same asymmetry the attacker exploits in Attack Scaling and the Sampling Budget, run in the defender's direction. Almost all traffic is benign, so the expected cost of a cascade is dominated by the cheap stage, while the expensive stage sets the strength of the defence on the small suspicious tail.
When it breaks
The red-team evidence is an absence, not a bound. The first system survived more than 3,000 hours of bug-bounty red teaming by 183 participants with no universal jailbreak found, and then a public demo drew about 3,700 hours from 339 participants and produced one, with four people clearing all eight levels. The defence did not change between those two results; the budget did.
The classifier is a model, so it has the attack surface of a model. Prompt-level evasion, distribution shift away from the synthetic training data, and poisoning of the classifier's own fine-tuning data are all live. Anthropic's own safeguards research has shown classifier fine-tuning datasets can be backdoored (Anthropic Alignment Science, 2026, Poisoning Fine-tuning Datasets of Constitutional Classifiers).
False refusals are not evenly distributed. A 0.05% aggregate refusal rate on harmless traffic can be a 5% rate for a medical researcher or a security engineer. Aggregate overrefusal metrics hide exactly the populations that notice.
Layers correlate. Defence in depth assumes independent failures. When the input classifier, output classifier and the model's own refusal training are all distilled from related data and related rules, an attack that defeats one is more likely than chance to defeat the others, and the composite is weaker than the product of the individual rates.
References and further reading
Every source this page cites, in the order it cites them. All of them open in a new tab.
- Sharma et al., 2025, Constitutional Classifiers, arXiv:2501.18837 arxiv.org
- Anthropic, 2025, Constitutional Classifiers anthropic.com
- Cunningham, Wei et al., 2026, Constitutional Classifiers++, arXiv:2601.04603 arxiv.org
- Anthropic, 2026, Next-generation Constitutional Classifiers anthropic.com
- Anthropic Alignment Science, 2026, Poisoning Fine-tuning Datasets of Constitutional Classifiers alignment.anthropic.com
7 flashcards for this concept
Click a card to reveal the answer.