Sparsity & Pruning intermediate 7 min read 6 flashcards

Activation Sparsity in Feedforward Layers

Trained transformers fire only a few percent of their MLP neurons on any given token, which is a property of the activations rather than the weights, and it is the reason a dense model can be computed sparsely without deleting anything.

Feed one input through T5-Base and count how many entries of the post-ReLU MLP activation are non-zero. The answer is about 3.0 percent. For ViT-B/16 it is 6.3 percent (Li et al., 2022, The Lazy Neuron Phenomenon: On Emergence of Activation Sparsity in Transformers, arXiv:2210.06313). Nobody pruned these models. Nobody added a sparsity penalty. The sparsity appeared on its own during ordinary training, and it gets stronger as models get wider and deeper.

That single observation separates this family of methods from everything in magnitude pruning and structured versus unstructured sparsity. Pruning deletes weights permanently and every token pays the same reduced cost. Activation sparsity deletes nothing: the weight matrix stays dense on disk, and a different subset of its rows turns out to be irrelevant for each token.

Where the zeros come from

A gated feedforward block computes, for hidden state \(x \in \mathbb{R}^{d}\),

\[\mathrm{FFN}(x) = W_{\text{down}}\big(\sigma(W_{\text{gate}}\,x) \odot W_{\text{up}}\,x\big)\]

with \(W_{\text{gate}}, W_{\text{up}} \in \mathbb{R}^{d_{\text{ff}} \times d}\) and \(W_{\text{down}} \in \mathbb{R}^{d \times d_{\text{ff}}}\). The intermediate vector has \(d_{\text{ff}}\) entries, typically three or four times \(d\). If entry \(j\) of that vector is zero, then row \(j\) of \(W_{\text{up}}\) and column \(j\) of \(W_{\text{down}}\) contributed nothing, and with ReLU the row of \(W_{\text{gate}}\) is the only part you actually needed to read.

So the unit of sparsity is a neuron: one row of the up projection paired with one column of the down projection. At 90 percent sparsity you are reading 10 percent of three large matrices rather than all of three large matrices. On a decode step that is a bandwidth saving, not a FLOP saving, and the distinction matters enormously (see cashing in activation sparsity on real hardware).

Modern models threw the zeros away

Llama, Mistral and Gemma use SwiGLU or GeGLU. SiLU and GELU are smooth and non-zero almost everywhere, so the exact-zero count is essentially nil. The sparsity did not disappear; the representation of it did. Define a neuron as inactive when its output magnitude falls below a threshold rather than when it is exactly zero, and SwiGLU models turn out to be sparse too, with a long tail of near-zero activations that contribute almost nothing to the output (Zhang et al., 2024, ReLU² Wins: Discovering Efficient Activation Functions for Sparse LLMs, arXiv:2402.03804). The practical form of that observation is that the distribution of hidden-state entries in Llama-architecture models is close to zero-mean and unimodal, which makes a magnitude threshold a well-behaved sparsifier (Liu et al., 2024, Training-Free Activation Sparsity in Large Language Models, arXiv:2408.14690).

This is why "how sparse is this model" is not a well-posed question until you say what counts as active. A threshold at \(10^{-3}\) and a threshold at \(10^{-1}\) give wildly different numbers for the same checkpoint, and papers choose thresholds that keep perplexity within a stated budget rather than reporting a raw zero count.

What makes a model sparser

The scaling behaviour has been measured directly. For decoder-only transformers, the activation ratio of ReLU models falls with training data as a decreasing logspace power law, while SiLU models trend the other way, becoming less sparse as training proceeds; at matched downstream performance the ReLU model is always the sparser one (Luo et al., 2024, Sparsing Law: Towards Large Language Models with Greater Activation Sparsity, arXiv:2411.02335). The same study finds that the activation ratio rises roughly linearly with the width-to-depth ratio up to a bottleneck point, which favours deeper, narrower models at a fixed parameter budget, and that the limiting sparsity is only weakly sensitive to parameter scale.

A later sweep across contemporary families reports that the sparsity a model can tolerate before quality moves tends to increase with size, and that the property survives instruction tuning and reasoning post-training (Haziza et al., 2025, Universal Properties of Activation Sparsity in Modern Large Language Models, arXiv:2509.00454). Bigger models are a better bet for this technique, not a worse one.

When it breaks

The measurement is the method. Every sparsity figure in this literature is relative to an activation definition and a quality budget. Compare two numbers only when both were produced with the same definition, and treat raw zero counts on a SwiGLU model as meaningless.

Attention is not the MLP. FFN neurons and attention heads are different sparsity substrates with different statistics, and a model can be 95 percent sparse in one and barely sparse in the other.

Prefill is dense. During prompt processing, hundreds or thousands of token positions run through the same weights at once, each selecting a different neuron subset. The union is close to everything, so the per-token sparsity buys nothing. This is a decode-time property.

Sparsity is not the same as skippability. A neuron whose activation is small still contributes. What the threshold methods exploit is that the sum of many small contributions is small, and that holds on average over a calibration distribution, not on every input.

References and further reading

Every source this page cites, in the order it cites them. All of them open in a new tab.

  1. Li et al., 2022, The Lazy Neuron Phenomenon: On Emergence of Activation Sparsity in Transformers, arXiv:2210.06313 arxiv.org
  2. Zhang et al., 2024, ReLU² Wins: Discovering Efficient Activation Functions for Sparse LLMs, arXiv:2402.03804 arxiv.org
  3. Liu et al., 2024, Training-Free Activation Sparsity in Large Language Models, arXiv:2408.14690 arxiv.org
  4. Luo et al., 2024, Sparsing Law: Towards Large Language Models with Greater Activation Sparsity, arXiv:2411.02335 arxiv.org
  5. Haziza et al., 2025, Universal Properties of Activation Sparsity in Modern Large Language Models, arXiv:2509.00454 arxiv.org
Check yourself

6 flashcards for this concept

Click a card to reveal the answer.

Drill the whole track