Model Architecture

Attention Sinks and Massive Activations: The Register the Transformer Had to Invent

A LLaMA-2-7B hidden state holds roughly 40,000 numbers. Four of them are more than ten thousand times larger than the median, and zeroing those four destroys the model. They exist because softmax attention has no way to say nothing here matters, so the network builds that missing operation out of the first token and pays for it in off-scale activations.

Take a trained LLaMA-2-7B, run a sentence through it, and look at the residual stream in the middle of the network. There are roughly 40,000 activations in that one hidden state. At most four of them are more than four orders of magnitude larger than the median value, and they sit at the same two feature dimensions, 1415 and 2533, on the same tokens: the start of the sequence, and the first period or newline. Set those four numbers to zero and the model collapses. Replace them with their mean values across inputs and the model is fine (Sun et al., 2024, Massive Activations in Large Language Models, arXiv:2402.17762).

That is not a bug, nor a quirk of one checkpoint. It is the visible residue of a missing primitive. Softmax attention must distribute exactly one unit of probability mass across the available keys at every position, so it has no representation for "this head has nothing to contribute here". Transformers solve that by electing a token every query can see, almost always the first, and dumping the unwanted mass onto it. The elected token becomes a register the architecture never provided, and forging it out of a real token is paid for in activations that run off the end of every numeric range the rest of the stack assumed.

Why this matters: Attention sinks are load-bearing. Almost every technique that makes inference cheaper touches them: sliding-window caches evict them, KV quantisation clips them, sparse attention skips them, and INT8 activation quantisation is wrecked by the outliers they create. If you do not know the register is there, you will delete it and watch perplexity explode for reasons that look inexplicable from the outside.

TL;DR

  • Softmax forces attention weights to sum to 1, so a head that wants to write nothing must route nearly all its mass to a token whose value vector is useless. Under causal masking the first token is the only one every query can always see, so it gets the job.
  • To hold \(1-\epsilon\) of the mass against \(n\) competing keys, the sink's logit must beat them by at least \(\ln(n/\epsilon)\). Logits are bounded by query and key norms, so a stronger sink demands larger norms. That is where massive activations come from, and it predicts that longer contexts and bigger models need stronger sinks.
  • Measured, not theorised: about 80% of attention heads in LLaMA 3.1 405B form strong sinks, and sink strength rises with model size and training context length (Barbero et al., 2025, arXiv:2504.02732).
  • Evicting the first tokens from a sliding KV cache does not lose information so much as fabricate it. Mass parked on a near-zero value vector renormalises onto whatever is left, turning a no-op head into a full-strength update of arbitrary content. Keeping four initial tokens restores stable generation out to 4 million tokens, up to 22.2x faster than sliding-window recomputation (Xiao et al., ICLR 2024, arXiv:2309.17453).
  • The same mechanism is why INT8 activation quantisation was hard: at the 6.7B scale a phase transition puts 150,000 outlier values per sequence into just 6 feature dimensions, in every layer (Dettmers et al., NeurIPS 2022, arXiv:2208.07339).
  • Vision transformers do the same thing with high-norm tokens in empty background patches, and the fix there was explicit register tokens (Darcet et al., ICLR 2024, arXiv:2309.16588).
  • The architectural repair is shipping. gpt-oss gives every head a learned scalar in the softmax denominator, so a head can attend to nothing without electing a victim token (OpenAI, 2025, gpt-oss-120b and gpt-oss-20b Model Card, arXiv:2508.10925). Replacing softmax outright also works: a rectified, non-sum-to-one variant drove the sink rate to 0% and cut hidden-state kurtosis from 33,510 to 340 in a 340M model (Zuhri et al., 2025, Softpick, arXiv:2504.20966).

At a Glance

The chain runs from a constraint to a workaround to a set of downstream failures. Every box after the first is something the architecture was not asked to do.

flowchart LR
  C["Softmax must sum to 1"] --> N["Head needs a no-op"]
  N --> S["Elect first token as sink"]
  S --> L["Logit gap must be large"]
  L --> M["Massive activations in residual"]
  M --> Q["INT8 quantisation breaks"]
  S --> E["Cache eviction breaks"]
  S --> F["Attention maps mislead"]
  M --> R["Fix: explicit null channel"]
  classDef slate fill:#334155,stroke:#64748b,stroke-width:1px,color:#e2e8f0
  classDef purple fill:#6d28d9,stroke:#a78bfa,stroke-width:1px,color:#fff
  classDef amber fill:#b45309,stroke:#fbbf24,stroke-width:1px,color:#fff
  classDef rose fill:#be123c,stroke:#fb7185,stroke-width:1px,color:#fff
  classDef emerald fill:#047857,stroke:#34d399,stroke-width:1px,color:#fff
  class C slate
  class N,S purple
  class L,M amber
  class Q,E,F rose
  class R emerald

Read it right to left: three well-known production failures share one upstream cause, and one architectural change removes that cause instead of patching each symptom.

[IMAGE: Heatmap of attention weights for one head of a mid-depth layer over a 64-token prompt, with the first column saturated and the rest near zero, beside a second heatmap from the same head with the first column removed and the remaining weights visibly redistributed. Caption: "The sink is not a feature of the input; it is a property of the head."]

Before Anyone Called It a Sink

The constraint arrived with the architecture. The original Transformer defines attention as a softmax over scaled dot products (Vaswani et al., 2017, Attention Is All You Need, arXiv:1706.03762), and softmax has no null output. Nobody noticed for five years because nobody was looking at activation magnitudes.

Quantisation made people look. Running 8-bit inference on OPT models, Dettmers and colleagues found that above roughly 6.7 billion parameters something sharp happens: outlier features stop being scattered and turn systematic, appearing in all layers and about 75% of sequence positions, concentrated in a handful of feature dimensions (Dettmers et al., NeurIPS 2022, arXiv:2208.07339). The fix was mechanical rather than explanatory: LLM.int8() isolates those dimensions into a 16-bit matrix multiply and does more than 99.9% of the arithmetic in 8-bit.

Two 2023 results identified the cause independently. Evan Miller, in a blog post with the best title in the literature, argued that the softmax denominator is off by one: a head forced to allocate its full mass must inflate its logits to approximate zero, so add a constant to the denominator and the pressure disappears (Miller, 2023, Attention Is Off By One). Bondarenko, Nagel and Blankevoort made the same argument with experiments. Outliers come from heads trying to learn a no-op or a partial residual update; the exact zeros needed in the attention matrix push the softmax input larger and larger during training; and two small modifications, clipped softmax and gated attention, let pre-trained BERT and OPT models reach full INT8 activation quantisation with no loss (Bondarenko et al., NeurIPS 2023, arXiv:2306.12929).

The vision side arrived at the same place that autumn from the opposite direction. ViT feature maps contain high-norm artifact tokens sitting in low-information background patches, which the network repurposes for internal computation; adding learnable register tokens removed the artifacts entirely and set a new state of the art on dense prediction (Darcet et al., ICLR 2024, arXiv:2309.16588). That paper named the thing correctly. The model needs scratch space, and if you do not give it any it will steal some.

StreamingLLM supplied the name that stuck and the practical consequence. Xiao, Tian, Chen, Han and Lewis, trying to make a model with a finite training window generate indefinitely under a sliding KV cache, found quality collapsed the moment the earliest tokens were evicted, and showed that retaining four initial tokens restores stable language modelling out to 4 million tokens with no fine-tuning (Xiao et al., ICLR 2024, arXiv:2309.17453).

[IMAGE: Timeline strip with two parallel lanes, "quantisation community" above and "long-context community" below, showing LLM.int8 and Quantizable Transformers on the upper lane and StreamingLLM and KVQuant on the lower, converging on a single merged lane in 2024 labelled Massive Activations. Caption: "Two communities, one phenomenon, three years apart."]

timeline
  title From an unexplained outlier to an architectural fix
  2017 : Transformer : softmax attention with no null option
  2022 : LLM.int8 finds systematic outliers : phase transition near 6.7B parameters
  2023 : Attention Is Off By One proposes softmax plus one
       : Quantizable Transformers ties outliers to no-op heads
       : StreamingLLM names the attention sink
       : Vision Transformers Need Registers finds the same artifact
  2024 : Massive Activations maps the outliers to fixed dimensions and tokens
       : KVQuant makes KV quantisation sink-aware
  2025 : Gu et al. show sinks act as key biases
       : Barbero et al. explain sinks as over-mixing control
       : gpt-oss ships learned per-head sink biases
       : Softpick removes sinks by replacing softmax

How a Transformer Builds a No-Op

Start from what a single head writes into the residual stream. For query position \(t\) with keys \(k_1 \dots k_n\) and values \(v_1 \dots v_n\), the head's output is a convex combination:

\[o_t = \sum_{i=1}^{n} a_i v_i, \qquad a_i = \frac{\exp(\ell_i)}{\sum_{j=1}^{n} \exp(\ell_j)}, \qquad \ell_i = \frac{q_t \cdot k_i}{\sqrt{d_k}}\]

The weights \(a_i\) are non-negative and sum to one. There is no setting of \(q_t\) that produces \(o_t = 0\) unless the value vectors themselves conspire to cancel. A head that has learned a narrow, conditional job, fire only when the previous token is a quotation mark, say, has to do something with its mass on the 99% of tokens where its job does not apply.

There are exactly two options. Either every value vector in reach is nearly zero, which cannot be arranged because the same value vectors are needed when the head does fire, or nearly all the mass goes to one position whose value vector is close to zero. The second is cheap, and it is what training finds. Gu and colleagues, sweeping optimisation, data distribution, loss and architecture across pre-training runs, report that the sink behaves less like attention and more like a key bias: it stores surplus attention score, is non-informative, and contributes almost nothing to the value computation (Gu et al., ICLR 2025, arXiv:2410.10781).

Why the register has to be loud

Now the quantitative part, which is the whole explanation of massive activations. Suppose position \(s\) is the sink and its logit exceeds every other logit by at least \(\Delta\). Then the leakage onto real tokens is bounded:

\[1 - a_s = \frac{\sum_{i \neq s} \exp(\ell_i - \ell_s)}{1 + \sum_{i \neq s} \exp(\ell_i - \ell_s)} \le n e^{-\Delta}\]

So holding all but \(\epsilon\) of the mass on the sink requires

\[\Delta \ge \ln \frac{n}{\epsilon}\]

Three consequences follow directly. First, the required logit gap grows with the logarithm of the context length, so a model trained on longer sequences needs a stronger sink. Second, \(\ell_i = q_t \cdot k_i / \sqrt{d_k}\) is bounded by \(\lVert q_t \rVert \lVert k_i \rVert / \sqrt{d_k}\), so the only way to buy a bigger gap is to grow the norms of the query or the sink's key. Third, those norms are produced from the residual stream by linear maps, which means the residual stream itself has to carry unusually large values at the sink position. The massive activations are not a side effect of the sink; they are the mechanism by which the sink is strong enough to be useful.

Barbero and colleagues reach the same place from a different direction. Deep transformers suffer from over-mixing: repeated attention layers smear token representations together until local distinctions are lost, the representational-collapse and over-squashing story familiar from graph neural networks. A sink is a controlled leak that lets a layer decline to mix, and their refined analysis predicts that bigger models and longer training contexts should show stronger sinks. Both predictions hold on the LLaMA 3.1 family, where roughly 80% of heads in the 405B model form strong sinks (Barbero et al., COLM 2025, arXiv:2504.02732).

The most recent synthesis ties the two literatures together with a proof: massive activations necessarily produce representational compression, with a bound on the entropy reduction. When the beginning-of-sequence token develops an extreme activation norm in the middle layers, attention sinks and "compression valleys", layers where representational entropy drops sharply, appear together. The authors verify this from 410M to 120B parameters and propose a "Mix-Compress-Refine" account of how depth is organised (Queipo-de-Llano et al., 2025, arXiv:2510.06477).

[IMAGE: Two-panel figure. Left: log-scale plot of maximum residual-stream activation magnitude by layer index for a 7B model, flat and small for the first two layers, jumping four orders of magnitude, holding constant through the middle, then falling in the last layers. Right: the same x-axis plotting mean sink attention mass per layer, showing the two curves rise and fall together. Caption: "The loud activation and the strong sink are one phenomenon measured two ways."]

What the sink is not

Three claims about the sink are wrong and all three appear in otherwise careful write-ups. It does not store information about the sequence: zeroing massive activations destroys the model, but fixing them to their dataset mean values preserves performance, the signature of a constant bias rather than a content-dependent signal (Sun et al., 2024). It is not a property of the first token's identity: sinks form regardless of how, or whether, a beginning-of-sequence token is included during pre-training, because what the role requires is visibility to every query (Barbero et al., 2025). And it is not present at initialisation; it emerges only after effective optimisation on sufficient data (Gu et al., ICLR 2025).

Seeing It in Motion

The failure that made all of this urgent is a cache policy, and it traces better as a sequence than a static picture. A naive sliding window keeps the most recent \(N\) tokens, and the sink is older than \(N\) almost immediately.

sequenceDiagram
  participant G as Generation loop
  participant C as KV cache
  participant H as Attention head
  G->>C: append token, evict oldest
  C-->>H: keys without sink
  H->>H: renormalise softmax over window
  Note over H: mass that parked on a near-zero value<br/>now lands on real values
  H-->>G: full-strength update, arbitrary content
  G->>G: next logits shift, perplexity rises
  Note over G,C: error compounds every step;<br/>keeping 4 initial tokens stops it

Three designs answer the same question, and only one removes the need for a victim token.

flowchart TB
  subgraph A["Implicit sink, status quo"]
    A1["First token key and value"] --> A2["Huge logit, tiny value"]
    A2 --> A3["Massive activations downstream"]
  end
  subgraph B["Preserve the sink, StreamingLLM"]
    B1["Pin first four KV entries"] --> B2["Slide window over the rest"]
    B2 --> B3["Stable to millions of tokens"]
  end
  subgraph D["Explicit null channel"]
    D1["Learned per-head scalar"] --> D2["Added to softmax denominator"]
    D2 --> D3["No victim token, no outlier"]
  end
  classDef amber fill:#b45309,stroke:#fbbf24,stroke-width:1px,color:#fff
  classDef teal fill:#0e7490,stroke:#22d3ee,stroke-width:1px,color:#fff
  classDef emerald fill:#047857,stroke:#34d399,stroke-width:1px,color:#fff
  class A1,A2,A3 amber
  class B1,B2,B3 teal
  class D1,D2,D3 emerald

The middle column needs no retraining, which is why it was adopted within weeks. The right column is a pre-training decision, which is why it took until 2025 to reach shipped weights.

[IMAGE: Line chart of perplexity against generated token position for three cache policies over 20,000 tokens: dense attention as a flat reference line, naive sliding window rising sharply the moment the window passes the first tokens, and sink-preserving window tracking the dense line. Caption: "The divergence point is the eviction of token zero, not the end of the training context."]

Watch It Run

Animated diagram showing attention mass flowing from a query to a sink token, the sink feeding a massive activation into the residual stream, and a feedback loop from the residual stream back into query and key norms.
Solid animated edges carry attention mass and residual-stream values in the direction of the arrow. The animated self-loop on the residual stream is the amplification cycle: large activations produce large key norms, which produce a larger logit gap, which concentrates more mass on the sink. The amber branch is what a quantiser sees; the green branch is the learned-bias alternative that bypasses the whole cycle. The Mermaid figures above show the same structure if the animation is unavailable.

By the Numbers

Quantity Value Model or setting Source
Massive activations above median more than 4 orders of magnitude LLaMA-2-7B, mid layers Sun et al., 2024
Count of massive activations at most 4 of roughly 40,000 per hidden state LLaMA-2-7B Sun et al., 2024
Where they sit dimensions 1415 and 2533; start token and first "." or newline LLaMA-2-7B Sun et al., 2024
Heads forming strong sinks about 80% LLaMA 3.1 405B Barbero et al., 2025
Outlier values per sequence about 150,000 in only 6 feature dimensions 6.7B transformer Dettmers et al., 2022
Scale of the outlier phase transition about 6.7B parameters OPT family Dettmers et al., 2022
Arithmetic kept in 8-bit by LLM.int8() more than 99.9% of values 6.7B and above Dettmers et al., 2022
Sink tokens needed for stable streaming 4, out to 4 million tokens Llama-2, MPT, Falcon, Pythia Xiao et al., 2024
Speedup over sliding-window recomputation up to 22.2x streaming setting Xiao et al., 2024
Hidden-state kurtosis, softmax to softpick 33,510 to 340 340M parameter model Zuhri et al., 2025
Sink rate with rectified softmax 0%, with 46.97% attention sparsity 340M and 1.8B models Zuhri et al., 2025
gpt-oss window attention bandwidth 128 tokens, alternating with dense layers gpt-oss-120b and 20b OpenAI, 2025
Models checked for the sink-compression link 410M to 120B parameters multiple families Queipo-de-Llano et al., 2025

Sources: Sun et al., 2024; Barbero et al., 2025; Dettmers et al., 2022; Xiao et al., 2024; Zuhri et al., 2025; the gpt-oss model card; Queipo-de-Llano et al., 2025. The dimension indices and the 22.2x speedup are specific to the checkpoints and hardware those papers measured; treat neither as a constant across model families.

A Concrete Example

Work one head, one position, by hand. The head sits in a mid-depth layer and its learned job is narrow, so on this token it has nothing to say. There are four keys in reach: the sink at position 0, and recent tokens at 7, 8 and 9. The scaled dot products come out as

sink (pos 0): 11.0
pos 7       :  2.1
pos 8       :  1.8
pos 9       :  2.4

Step 1, exponentiate relative to the maximum. Subtracting 11.0 from everything leaves \(e^0 = 1\) for the sink and three small numbers: \(e^{-8.9} = 1.36 \times 10^{-4}\), \(e^{-9.2} = 1.01 \times 10^{-4}\), \(e^{-8.6} = 1.83 \times 10^{-4}\).

Step 2, normalise. The three non-sink terms total \(4.20 \times 10^{-4}\), so the denominator is \(1.000420\) and the weights are

a_sink = 0.999580
a_7    = 0.000136
a_8    = 0.000101
a_9    = 0.000183

Step 3, compute the head output. Suppose the sink's value vector has norm 0.05, which is what "the model learned this value is useless" looks like numerically, and the three real value vectors have norms near 1.2. The output norm is about \(0.99958 \times 0.05 + 4.20 \times 10^{-4} \times 1.2 \approx 0.0505\). Against a residual stream whose norm at this layer is on the order of 60, the head has written roughly 0.08% of the signal. It has successfully done nothing, which is what it wanted.

Step 4, evict the sink. Run the same step with a sliding window that has dropped position 0. The three remaining exponentials are unchanged, but they are now the whole denominator:

a_7 = 1.36 / 4.20 = 0.324
a_8 = 1.01 / 4.20 = 0.240
a_9 = 1.83 / 4.20 = 0.436

The output is a convex combination of three real value vectors with norm about 1.2, so the head now writes roughly 1.2 instead of 0.05: a 24x larger update, carrying content the head never intended to select. Nothing was forgotten here. Something was fabricated. Multiply that across dozens of heads and dozens of layers, feed the corrupted logits back in as the next token, and the compounding is immediate, which is exactly the perplexity explosion Xiao and colleagues measured.

Step 5, the architectural fix, same arithmetic. Give the head a learned scalar \(s\) in the denominator and let it reach 11.0 during training:

\[a_i = \frac{\exp(\ell_i)}{\exp(s) + \sum_j \exp(\ell_j)}\]

With \(s = 11.0\) and the same three logits, \(0.99958\) of the mass goes to the \(\exp(s)\) term, which multiplies no value vector at all. The output norm is \(4.20 \times 10^{-4} \times 1.2 \approx 0.0005\), a cleaner no-op than the sink achieved, with no KV entry to protect, no value vector forced toward zero, and no pressure on any key norm. This is the gpt-oss design: a learned per-head bias in the softmax denominator, not a token in the sequence (OpenAI, 2025, arXiv:2508.10925).

[IMAGE: Four-bar grouped chart of attention weight on a log y-axis for the three scenarios in the worked example: with sink, sink evicted, and learned bias, with the resulting head output norm annotated above each group. Caption: "Same logits, three denominators, two orders of magnitude of difference in what the head writes."]

Where It Breaks

KV cache eviction and compression

Any policy that drops or compresses old entries has to special-case the sink: sliding windows, heavy-hitter eviction, page-level selection, and any scheme that compresses by clustering keys. The rule from StreamingLLM is to pin the first few entries unconditionally. The trap is that the number is empirical; four sufficed for the models Xiao and colleagues tested, and treating four as a constant across architectures and training recipes is an approximation.

Quantisation, both weights and KV

This is where the damage is quietest. A per-tensor activation scale computed over a distribution containing values four orders of magnitude above the median spends its entire dynamic range on four numbers, and everything else collapses toward zero. That is why INT8 activation quantisation needed mixed-precision decomposition at all. The KV cache has its own version: keys carry outliers in specific channels while values do not, so per-channel for keys and per-token for values is the right asymmetry, and the first token still needs separate treatment. KVQuant is explicit about sink-aware quantisation for this reason (Hooper et al., NeurIPS 2024, arXiv:2401.18079).

Sparse and windowed attention

Trainable sparse attention and banded windows both risk masking the sink out of reach for some queries. gpt-oss is instructive: it alternates banded 128-token windows with dense layers and gives every head a learned sink bias, so a windowed layer never depends on a token outside its band. A sparse design that keeps the implicit sink without guaranteeing every query can reach it is a trap for whichever head learned to rely on it.

Reading attention maps as explanations

An attention map where one column holds 99.9% of the mass is not telling you what the model attended to. It is telling you the head declined to attend. Any analysis that ranks token importance by raw attention weight, or visualises a head without separating the sink column, will keep reporting that models care enormously about the beginning-of-sequence token. Methods that intervene on activations rather than read attention weights avoid this, which is part of why the field moved toward causal intervention; see circuit tracing and activation patching.

Pooling and embeddings

Mean-pooling hidden states into a sentence embedding averages in a vector with one coordinate four orders of magnitude off scale. The pooled vector is dominated by the sink dimension, cosine similarities compress toward one, and the space looks anisotropic for reasons unrelated to semantics. Excluding the first token and the first delimiter, or pooling a final-layer representation where massive activations have already diminished, is a cheap mitigation. See embedding space geometry and anisotropy.

Template changes and context extension

Because the sink is positional rather than semantic, anything that shifts what sits at position zero perturbs it: a new chat template, dropping the beginning-of-sequence token during fine-tuning, concatenating documents without separators, or serving from a prefix cache that begins mid-document. Models usually recover, since they will elect whatever is at position zero, but the transition costs quality and the symptom looks like an unrelated regression. Context extension is the sharper version of the same problem: position-interpolation methods raise \(n\) without changing the key norms that set \(\Delta\), so leakage onto real tokens grows. That is one of several mechanisms behind degradation in naively extended models, alongside the positional effects covered in long-context RoPE scaling.

Alternative Designs

Design How it works Key advantage Key limitation Best when
Implicit sink, no intervention Training elects position 0 and inflates its logit Nothing to implement; it is the default Massive activations; every efficiency technique must special-case it You are consuming existing weights and only need to avoid breaking them
Pin the first k KV entries Keep 4 initial entries, slide the window over the rest Inference-time only, no retraining, stable to millions of tokens Does not recover evicted content; k is empirical Streaming or long-running generation on an existing checkpoint
Dedicated sink token in pre-training Prepend a learnable placeholder token to every sequence A single pinned entry suffices; sink behaviour becomes predictable Requires pre-training; still a token with a value vector You control pre-training and want minimal architectural change
Learned per-head sink bias Add a learned scalar to the softmax denominator per head No victim token, no KV entry, head can attend to nothing Pre-training decision; existing kernels need the extra term Designing a new model, especially with windowed attention
Softmax plus one Add a constant 1 to the denominator One-line change, removes the pressure to inflate logits Not learned, so the null capacity is fixed across heads and layers Research and ablation; a cheap first experiment
Clipped softmax and gated attention Allow exact zeros, or gate the head's output path Full INT8 activation quantisation with no quality loss, demonstrated on BERT and OPT Changes the attention module; pre-training required Quantisation is the primary objective
Rectified non-sum-to-one softmax Replace softmax with a variant that need not sum to 1 0% sink rate, kurtosis down two orders of magnitude, sparse attention maps New operator, kernel support immature, mainly validated below 2B parameters so far Greenfield pre-training where low-bit serving is the target
Register tokens Append learnable tokens with no input dependence Fixes high-norm artifacts; better dense prediction features Extra sequence length; established for ViTs rather than causal LMs Vision backbones and dense feature extraction

Nothing here is free. Everything below the second row requires a pre-training commitment, and the one option that does not, pinning initial KV entries, treats the symptom. The field spent two years learning to protect an accident and is now, in new models, designing it away.

How It Is Used in Practice

Sink preservation is standard in serving stacks. Any implementation of sliding-window or streaming attention that works does something equivalent to pinning initial entries, and quantisation toolchains either isolate outlier channels or apply a transform that spreads them before rounding. Those two behaviours look unrelated in a configuration file and are the same accommodation.

The architectural turn is visible in shipped weights. gpt-oss uses grouped-query attention with learned attention-sink biases in the softmax denominator, 64 query heads of dimension 64 against 8 key-value heads, alternating banded 128-token windows with dense layers (OpenAI, 2025, arXiv:2508.10925). The learned bias is what makes the aggressive windowing safe: a head in a banded layer can decline to attend without needing a token that may lie outside its band. The sink bias is not a tweak beside the efficiency choice, it is its enabling condition.

The operational advice is short. Compressing a KV cache: protect the sink positions and verify perplexity past the window boundary, not only inside it. Quantising activations: measure per-channel kurtosis before choosing a scheme, because a per-tensor scale on a distribution with four extreme values fails in a way aggregate benchmarks hide. Pooling hidden states: exclude the first token and the first delimiter. Reading an attention map: subtract the sink column first. And when quality drops after a template or separator change, check whether you moved what sits at position zero.

[IMAGE: Annotated screenshot-style schematic of a serving configuration showing three settings that are usually tuned independently, sliding window size, KV quantisation scheme, and pinned sink count, with arrows connecting all three to a single shared cause. Caption: "Three knobs, one phenomenon."]

Insights Worth Remembering

  1. Softmax has no null output, and that absence is an architectural decision nobody made deliberately. Everything here follows from the sum-to-one constraint. A head that must put its mass somewhere finds the cheapest somewhere, and the cost shows up thousands of dimensions away.

  2. Massive activations are the price of a strong sink, not an independent pathology. The gap needed to hold \(1-\epsilon\) of the mass against \(n\) keys is at least \(\ln(n/\epsilon)\), logits are bounded by query and key norms, and those norms come from the residual stream. The chain is mechanical.

  3. Sink eviction fabricates content rather than losing it. Renormalisation moves mass parked on a near-useless value vector onto real ones, turning a no-op into a confident wrong update. That is why the failure is catastrophic rather than gradual, and why "the model forgot the start of the prompt" is the wrong diagnosis.

  4. The quantisation problem and the streaming problem are the same problem. Different communities found them a year apart, wrote them up in different vocabularies, and patched them separately. Treating them as separate means paying for both.

  5. An attention weight of 0.999 on token zero means the head is idle. Reading it as importance inverts the finding, and every importance measure built on raw attention weights inherits the error.

  6. The sink is positional, not semantic. It forms whether or not a beginning-of-sequence token exists, because the role requires visibility to every query rather than content. That is why the fix is a channel and not a better special token.

  7. Bigger and longer-context models need stronger sinks, and this was predicted before it was measured. The over-mixing analysis implied it, the LLaMA 3.1 family confirmed it at 405B, and the \(\ln n\) bound explains why context length specifically matters. Vision transformers independently invented the same workaround, and the register-token fix landed there first. When two modalities converge on the same hack, the hack is telling you something about the operator rather than the data.

Open Questions

Does removing sinks change what the model can learn, or only how it stores it? Known: a rectified non-sum-to-one softmax reaches 0% sink rate with comparable benchmark performance and better low-bit behaviour at 340M and 1.8B parameters (Zuhri et al., 2025). Unknown: whether that holds at frontier scale, where the over-mixing pressure the sink relieves is strongest. A follow-up has already argued that the rectified variant has an initialisation problem, so this is actively contested rather than settled.

How many sinks does a model actually need? Measured: four initial tokens suffice for the models StreamingLLM tested, and massive activations concentrate on at most four values per hidden state. Not established: whether the right number is a function of depth, head count, context length, or training data, or whether it is better thought of as total null capacity per layer rather than a count of positions.

Is the sink one mechanism or several wearing one name? The key-bias account, the over-mixing account and the compression account all fit the evidence and all three have supporting experiments. They may be the same claim in three vocabularies, or there may be distinct phenomena, for instance a true no-op register and a separate depth-wise compression device, that current measurements cannot separate.

Can a trained model's sink be retired after the fact? Open. Replacing massive activations with mean values preserves performance, which suggests the content is constant and therefore in principle absorbable into a bias term. Whether post-hoc surgery that moves the register into an explicit channel, and then permits aggressive low-bit activation quantisation, works without retraining is untested in published work. The same gap applies to tooling already deployed: raw attention weights are known to be contaminated, but nobody appears to have measured how much interpretability work with a token-importance step inherits the artifact.

Sources and Further Reading

  1. Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). "Attention Is All You Need." NeurIPS 2017. arXiv:1706.03762
  2. Dettmers, T., Lewis, M., Belkada, Y., & Zettlemoyer, L. (2022). "LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale." NeurIPS 2022. arXiv:2208.07339
  3. Miller, E. (2023). "Attention Is Off By One." evanmiller.org
  4. Bondarenko, Y., Nagel, M., & Blankevoort, T. (2023). "Quantizable Transformers: Removing Outliers by Helping Attention Heads Do Nothing." NeurIPS 2023. arXiv:2306.12929
  5. Darcet, T., Oquab, M., Mairal, J., & Bojanowski, P. (2024). "Vision Transformers Need Registers." ICLR 2024 (Outstanding Paper). arXiv:2309.16588
  6. Xiao, G., Tian, Y., Chen, B., Han, S., & Lewis, M. (2024). "Efficient Streaming Language Models with Attention Sinks." ICLR 2024. arXiv:2309.17453
  7. Sun, M., Chen, X., Kolter, J. Z., & Liu, Z. (2024). "Massive Activations in Large Language Models." COLM 2024. arXiv:2402.17762
  8. Hooper, C., Kim, S., Mohammadzadeh, H., Mahoney, M. W., Shao, Y. S., Keutzer, K., & Gholami, A. (2024). "KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization." NeurIPS 2024. arXiv:2401.18079
  9. Gu, X., Pang, T., Du, C., Liu, Q., Zhang, F., Du, C., Wang, Y., & Lin, M. (2025). "When Attention Sink Emerges in Language Models: An Empirical View." ICLR 2025 (Spotlight). arXiv:2410.10781
  10. Barbero, F., Arroyo, Á., Gu, X., Perivolaropoulos, C., Bronstein, M., Veličković, P., & Pascanu, R. (2025). "Why do LLMs attend to the first token?" COLM 2025. arXiv:2504.02732
  11. Zuhri, Z. M. K., Fuadi, E. H., & Aji, A. F. (2025). "Softpick: No Attention Sink, No Massive Activations with Rectified Softmax." arXiv:2504.20966
  12. OpenAI (2025). "gpt-oss-120b and gpt-oss-20b Model Card." arXiv:2508.10925
  13. Queipo-de-Llano, E., Arroyo, Á., Barbero, F., Dong, X., Bronstein, M., LeCun, Y., & Shwartz-Ziv, R. (2025). "Attention Sinks and Compression Valleys in LLMs are Two Sides of the Same Coin." arXiv:2510.06477
  14. Related concepts on this site: attention sinks, activation outliers and rotation quantisation, KV cache eviction and compression, natively trainable sparse attention, quantisation grids, scale and zero point, embedding space geometry and anisotropy.

Free to read, no ads, no sign-up. If it was useful you can buy me a coffee.