Safety & Alignment 1 October 2026 2 min read 464 words

Refusal was never in the weights

Anthropic's evaluation of GLM-5.3 lists three ways past the model's refusals. The free one gets 92% and the one costing thousands of dollars gets 100%. That eight-point gap says where a frontier safeguard actually sits, and what an open-weight release takes away.

The argument

Refusal under a prefilled reasoning trace is enforced by a server checking a signature, not by the weights, so an open-weight release deletes that safeguard rather than weakening it.

Three ways past the same model's refusals, measured in the same simulated environment, fifty samples per cell. Tell GLM-5.3 that it is an autonomous red-team agent on an authorised exercise, and it engages with an overtly malicious order against critical systems 64% of the time. Prefill its thinking tokens so that it appears to have already weighed the request and decided to proceed, and that rises to 92%. Download the weights and edit the refusal out of them, which cost Anthropic's team about 2,200 GPU hours and roughly $4,400, and it engages in every trial. Out of the box, with none of this, the model refused in all of them.

So the expensive attack beats the free one by eight percentage points. That ranking is the most interesting number in Anthropic's write-up of GLM-5.3, published on 29 September, and it is not the number anyone quoted. Weight surgery is the part of open-weight release that gets argued about, and on this evidence it is the part that barely matters. What matters is that writing the model's own reasoning for it works nearly as well, costs nothing, needs no GPUs, and is available to anyone who runs the model themselves.

Which tells you where the safeguard against it lives in the models that resist it. Not in the weights. In a signature.

Why a thought you wrote yourself is the strongest prompt

A transformer computes a distribution over the next token given every token before it. That is the whole interface. The sequence is a list of embeddings plus positions; there is no field anywhere in it that records who supplied which token. The boundary between what the user said and what the model said is itself made of tokens, drawn by the chat template the serving harness applies. You can read one in the open: the Llama-3 template in the refusal_direction repository wraps the instruction in <|start_header_id|>user<|end_header_id|>, closes it, and then writes <|start_header_id|>assistant<|end_header_id|> before handing control to the model. The harness writes the assistant header. Whatever follows it is, as far as the forward pass is concerned, the assistant speaking.

Refusal is a behaviour learned on top of that. Post-training teaches a model to produce a refusal when the context looks like a harmful request, and reasoning post-training additionally rewards answers that follow from the reasoning that preceded them. Those two lessons point in opposite directions the moment someone else gets to author the reasoning. A context in which the model has apparently just deliberated and concluded that it will proceed is not a request to be refused; it is a commitment to be honoured. The 64% from the cover story is the model conditioning on a claim about the world. The 92% from the prefill is the model conditioning on a claim about itself, and it is stronger because consistency is exactly what it was trained for.

This is not a bug a better training run obviously fixes, because the model has no way to tell the difference. Self-authorship is not represented in the input.

What abliteration actually does, and what it buys

The expensive route is worth understanding too, because it is usually described as something more mysterious than it is. The technique Anthropic used, and that several developers applied to GLM-5.3 publicly within days of its release, comes from the work on refusal being mediated by a single direction. Its code reads clearly. Run a batch of harmful instructions and a batch of harmless ones through the model, cache the mean activation at the end-of-instruction positions for every layer, and subtract: mean_diff = mean_harmful - mean_harmless. Each layer yields a candidate refusal direction \( \mathbf{r} \) in the residual stream. Pick the one that works best, normalise it, and then remove that component from the activation wherever it appears:

\[ \mathbf{a} \leftarrow \mathbf{a} - (\mathbf{a}\cdot\hat{\mathbf{r}})\,\hat{\mathbf{r}} \]

To make it permanent, the same projection is applied to every weight matrix that writes into the residual stream: the embedding matrix, each attention output projection, each MLP down-projection. Nothing in the network can write that direction any more. No retraining, no data, no gradient steps.

That is why capability survives. On GPQA-Diamond, Anthropic reports the abliterated and standard models scoring the same; on a subset of CyberGym, the abliterated version came in a few percent lower. The edit took GLM-5.3's refusal rate on three public harm benchmarks from above 90% to about 3%, 2% and 12%. Anthropic put its own cost at roughly $4,400 for a team that had never done it before, and estimated an experienced team would need closer to 600 GPU hours, about $1,200. The smaller GLM-5.3-Flash took about 600 GPU hours.

So abliteration is not pointless. It buys a model that never refuses anything, once, for everyone, with no per-request trickery. In the cyber scenario it bought eight points. For an attacker at scale, the convenience is worth more than the points.

The frontier's answer is cryptographic, not behavioural

Now look at what a hosted model does with the prefill attack. It does not resist it. It forbids it. Anthropic's API documentation is one line: "You can't prefill the assistant response while thinking is on."

And the mechanism behind that line is a provenance check on reasoning. Every thinking block comes back with a signature field, described in the docs as an encrypted copy of the full reasoning, which the API uses "to verify that thinking blocks were generated by Claude when you pass them back". Modified thinking blocks are rejected with a 400 error. From Claude Fable 5.1 onward the check goes further: the signature is verified against the model that produced the block, and against everything sent before it, so a block stays valid only while the system prompt, the tools and the preceding messages are unchanged. Edit the prefix and the block is invalid. Claude Sonnet 5.5's blocks are additionally tied to the account that produced them.

Read that as a security design and it is unambiguous. The frontier's defence against a prefilled reasoning trace is not a weight that holds under pressure. It is a server that will not accept a thought it did not write. It is the same philosophy Anthropic described in May for agent containment, where the stated aim was to supervise what an agent is able to do rather than what it does.

Which is why Anthropic's own comparison table has padlocks in it. Against safeguarded Claude models, the cover story was tested and failed. The prefill and the abliteration cells are marked as attacks "not generally feasible against the Claude API", because the API offers no way to prefill Claude's thinking and the weights are not available to edit. That is honest reporting, and it also means something precise: two of the three rows in the Claude column are facts about an interface, not measurements of a model. Nobody outside Anthropic knows how Claude's weights behave when the reasoning has been written for them, and nobody can find out.

The thing that cannot be shipped

Here is the consequence I think is being missed. The open-weight safety debate is usually a debate about training: did the lab post-train refusals properly, did it red-team enough, did it ship with safeguards. GLM-5.3 largely passes that test. It refuses above 90% on three public harm benchmarks. On a bare order to attack a remote system it refused in all fifty trials. Zhipu did train refusals.

What it cannot ship is the verifier. A signature on a reasoning trace is not a property of a function; it is a property of a relationship in which one party holds a key the other does not. Download the weights and you get the sampler, the chat template and the context assembly path along with them. There is no server left to tell you that a thinking block is forged, because you are the server. Any safeguard whose enforcement depends on knowing who wrote a token is therefore not transferable to an open-weight release, however good the post-training is. That class of defence does not degrade on the way out of the lab. It disappears.

The honest counterargument is that this distinction is academic. The threat model asks what an attacker can do, and the answer is the same either way: GLM-5.3-Flash turned a published Chrome CVE into a working ARM64 exploit chain, bypassing pointer-authentication hardening, on 20 minutes of human attention and eight hours of compute costing $20.40 at Zhipu's own prices. Nobody holding that capability cares which layer of the stack failed to stop them.

But the layer decides what anyone can usefully ask for. If the gap between GLM-5.3 and a safeguarded frontier model were a training gap, the remedy would be pressure on open-weight labs to train harder, and it would work. If the gap is an enforcement gap, that pressure cannot work, and the only remaining levers are the ones Anthropic actually reaches for at the end of its post: independent government evaluation of capability before release, decisions about what gets released at all, and getting equivalent capability to defenders. Those are unsatisfying asks. They are the right asks, and the eight-point gap is why.

Two cautions on the evidence. The engagement percentages come from a simulated world with a fake shell, where another model approximates what each command would have returned; Anthropic says plainly that this is an imperfect measure of real behaviour. And most of this record is one company's: Anthropic evaluating a competitor's open model, and documenting its own API. The NIST assessment it cites, which found GLM-5.3 the most cyber-capable open-weight model so far and about four months behind the US frontier, is quoted here at second hand.

For anyone learning this field, the habit worth forming is small. When you read a refusal or jailbreak-robustness number, ask what the attacker was permitted to write. A score measured through an interface that forbids prefills, signs reasoning, binds it to a prefix and never hands over weights is a measurement of a deployment. It tells you much less than it looks like it does about the model inside. And if you serve open weights yourself, benignly, you have taken the verifier's job: your context assembly path is now a safety boundary, which matters most in agents, where much of the context is written by tool output rather than by you.

There is a way to give an open-weight model the safeguard that holds the prefill attack at zero. You stop releasing the weights.

What this is argued from

Reporting and primary material the piece rests on, dated at the time of writing. The interpretation is mine; the facts belong to these.

  1. GLM-5.3 and the spread of advanced cyber capabilities Anthropic · 2026-09-29
  2. Thinking, including thinking encryption and preserved thinking Anthropic · 2026-10-01
  3. Extended thinking, manual mode and its limits Anthropic · 2026-10-01
  4. refusal_direction, code accompanying "Refusal in Language Models Is Mediated by a Single Direction" GitHub · 2026-10-01
  5. generate_directions.py, the difference-in-means extraction GitHub · 2026-10-01
  6. hook_utils.py, the directional ablation hooks GitHub · 2026-10-01
  7. llama3_model.py, weight orthogonalisation and the chat template GitHub · 2026-10-01
  8. How we contain Claude across products Anthropic · 2026-05-25

Editorials on this site are written to be argued with. If you think the reading is wrong, it probably is in some particular way, and that is the useful part.

refusalabliterationopen weightsreasoning tracesprefill