Inference & Serving

Activation Sparsity: The 90 Percent of a Dense Model That Does Nothing Per Token

Feed one token through T5-Base and 97 percent of its MLP neurons output exactly zero. Nobody pruned the model; the sparsity arrived on its own. The honest exchange rate for cashing it in is a 2.5x cut in arithmetic for 1.40x on a GPU. Here is why that gap exists, and where the technique still wins outright.

Push one input through T5-Base, stop at the feedforward block, and count the non-zero entries in the post-ReLU activation: about 3.0 percent. For ViT-B/16 it is 6.3 percent (Li et al., 2022, The Lazy Neuron Phenomenon: On Emergence of Activation Sparsity in Transformers, arXiv:2210.06313). Nobody pruned these networks and nobody added a sparsity penalty. The zeros showed up during ordinary training, in vision and in language, at every depth, and they get more extreme as models get wider.

The inference consequence writes itself. If 97 percent of a neuron bank contributes nothing to this token's output, 97 percent of two large matrices did not need to be read, and a 70-billion-parameter model would behave, for one token, like a much smaller one. The results of chasing that are real and much smaller than the arithmetic suggests: the most carefully measured recent figure is a 2.5x reduction in FLOPs converting to 1.40x of wall-clock decoding on a GPU (Zhang et al., 2025, Spark Transformer: Reactivating Sparsity in FFN and Attention, arXiv:2506.06644). That ratio, not the sparsity percentage, is the subject.

Why this matters: Activation sparsity is the one efficiency technique that needs no retraining, no accuracy sacrifice and no architecture change, and the one whose benefit collapses fastest under the conditions real serving systems run in. Knowing why 90 percent sparsity becomes 15 percent latency tells you when to reach for it, when to reach for quantisation or Mixture of Experts instead, and why every success story here is a batch-size-one workload with a memory cliff in it.

TL;DR

  • Trained transformers fire a few percent of their MLP neurons per token with no intervention. ReLU models get sparser as they train; SiLU and GELU models get less sparse, a cost of the 2020 activation-function switch that nobody priced (Luo et al., 2024, arXiv:2411.02335).
  • The hard part is ordering: knowing a neuron is dead requires computing it. Every method here finds out before paying, with a learned predictor, a threshold on the layer's input, or a static hot-cold partition.
  • Deja Vu silenced over 80 percent of attention heads and over 95 percent of MLP parameters per token on OPT-175B, accuracy flat to roughly 75 percent sparsity, for over 2x lower latency than FasterTransformer (Liu et al., 2023, arXiv:2310.17157).
  • Nominal sparsity converts to elapsed time at a terrible rate: 50 percent sparsity bought CATS about 15 percent faster generation, and a 2.5x FLOP cut bought Spark Transformer 1.40x on GPU. Decode is bandwidth-bound, gathers are not free, and tensor cores want dense tiles.
  • Batching is the killer. At 90 percent per-token sparsity and a batch of 32, the union of active neurons covers roughly 73 percent of the layer with strong hot-neuron correlation and about 97 percent without it. Serving systems batch in order to escape the bandwidth bound, which is what destroys the sparsity.
  • As of early 2026 the sparsity that survives batching has moved from MLP neurons to attention heads, which stay selective because each request attends to its own context (Shrestha et al., 2025, arXiv:2505.14884).

At a Glance

flowchart LR
    A["Hidden state x"] --> P["Predictor or threshold"]
    P --> I["Active index list"]
    I --> G["Gather rows of W_up"]
    G --> M["Small dense GEMV"]
    M --> S["Scatter into output"]
    S --> O["Layer output"]
    A -.->|skipped entirely| D["90 percent of W_up and W_down"]
    B["Batch of 32 tokens"] --> U["Union of active sets"]
    U --> C["73 percent dense: saving gone"]

    classDef blue fill:#1e40af,stroke:#3b82f6,stroke-width:1px,color:#fff
    classDef purple fill:#6d28d9,stroke:#a78bfa,stroke-width:1px,color:#fff
    classDef teal fill:#0e7490,stroke:#22d3ee,stroke-width:1px,color:#fff
    classDef emerald fill:#047857,stroke:#34d399,stroke-width:1px,color:#fff
    classDef rose fill:#be123c,stroke:#fb7185,stroke-width:1px,color:#fff
    class A,B blue
    class P,I,G,M,S purple
    class O teal
    class D emerald
    class U,C rose

The top row is the mechanism; the bottom row is why most deployments never see it.

Before the Zeros Went Away

The original transformer feedforward block was two matrices with a ReLU between them. ReLU emits exact zeros for every negative pre-activation, so the sparsity sat in plain sight in every checkpoint from 2017 onwards and almost nobody exploited it, because training-time efficiency was the bottleneck everyone cared about.

Then two things happened in opposite directions. The 2022 measurement work established that the sparsity was not an artefact of small models or particular tasks: it strengthened with scale, which made it strategically interesting. Meanwhile the architecture community adopted GELU, SiLU and gated variants like SwiGLU on the strength of consistent small quality gains. Those functions are smooth and non-zero almost everywhere, so the exact-zero count in a Llama checkpoint is essentially nil. By 2023 the field had a property it wanted and models that no longer exhibited it in exploitable form, and everything since has been an attempt to resolve that by prediction, by thresholding, or by putting ReLU back.

timeline
    title From Accident to Engineering Target
    2017 : Transformer FFN uses ReLU, so exact zeros are free but unexploited
    2022 : Li et al. document the lazy neuron phenomenon
         : T5-Base fires 3.0 percent of MLP entries, ViT-B16 6.3 percent, and larger is sparser
    2023 : Deja Vu predicts contextual sparsity per token and cuts OPT-175B latency over 2x
         : ReLU Strikes Back argues for reinstating ReLU in LLMs
         : PowerInfer splits hot and cold neurons across GPU and CPU
         : LLM in a flash runs models twice the size of DRAM off flash storage
    2024 : ProSparse regularises LLaMA2-7B to 89.32 percent sparsity at comparable quality
         : TurboSparse dReLU reaches 90 percent while matching SwiGLU quality
         : CATS and TEAL obtain 40 to 50 percent with no retraining at all
         : Q-Sparse makes the sparsity level a top-K hyperparameter
    2025 : Polar Sparsity shows MLP sparsity vanishes under batching, moving the win to attention heads
         : Spark Transformer pretrains 8 percent FFN activation into a Gemma-2 recipe

How Activation Sparsity Actually Works

The unit of sparsity is a neuron, not a weight

A gated feedforward block computes, for a hidden state \(x \in \mathbb{R}^{d}\),

\[\mathrm{FFN}(x) = W_{\text{down}}\big(\sigma(W_{\text{gate}}\,x) \odot W_{\text{up}}\,x\big)\]

with \(W_{\text{gate}}, W_{\text{up}} \in \mathbb{R}^{d_{\text{ff}} \times d}\) and \(W_{\text{down}} \in \mathbb{R}^{d \times d_{\text{ff}}}\). The intermediate vector has \(d_{\text{ff}}\) entries, three to four times \(d\) in practice. If entry \(j\) of that vector is zero, then row \(j\) of \(W_{\text{up}}\) multiplied something that was discarded, and column \(j\) of \(W_{\text{down}}\) multiplied a zero. Neither needed to be in memory.

So a "neuron" here means a matched pair: one row of the up projection, one column of the down projection. That pairing is what makes the sparsity structured enough to be worth something, because removing whole rows and columns is a gather rather than a sparse format, and sparse formats rarely beat dense matmul on a GPU.

What it does not do is make the model smaller. Every weight is still needed, because some other token will need it. This is the exact inverse of pruning: storage unchanged, per-token read cost down.

[IMAGE: Two heatmaps of the same 128-by-128 slice of an FFN intermediate activation for two tokens from one prompt, white for zero and blue for active, row indices aligned. Caption: "Same weights, same layer, two tokens. The active sets overlap without coinciding, which is the premise and the problem."]

Why nobody can simply skip the dead neurons

The ordering problem in one sentence: to learn that neuron \(j\) is inactive you evaluate \(\sigma(W_{\text{gate}}[j,:]\,x)\), which requires reading the row you hoped to skip. Three escapes exist, and the field is a choice among them.

Predict it. Train a small network per layer that maps the layer input to a score per neuron, and keep the top scores. The claim underneath is that contextual sparsity is not merely present but predictable from the layer input by something much cheaper than the layer. Deja Vu made this concrete with a two-layer predictor on OPT-175B, where roughly 85 percent total sparsity was available because the model carries about twice as many MLP as attention parameters.

Threshold the input instead of the output. Sparsify \(x\) itself rather than the intermediate and the zeros remove columns from every projection that consumes \(x\), the gate included. No prediction needed, because what you threshold is already in registers. TEAL does this with a per-layer magnitude threshold on hidden states throughout the model, reaching 40 to 50 percent model-wide sparsity with minimal degradation across Llama-2, Llama-3 and Mistral from 7B to 70B (Liu et al., 2024, Training-Free Activation Sparsity in Large Language Models, arXiv:2408.14690).

Decide statically, once. Neuron activation frequency follows a power law: a small pool fires for nearly every token, a long tail is input-specific. PowerInfer pins the hot pool in GPU memory and computes cold neurons on the CPU, with no per-token decision for the hot part (Song et al., 2024, PowerInfer, SOSP 2024, arXiv:2312.12456).

The predictor's budget is exactly the sparsity it claims

A predictor is only worth running if it is cheaper than the work it avoids. Write \(B_{\text{dense}}\) for the bytes the dense layer would read, \(s\) for the predicted sparsity, and \(B_{\text{pred}}\) for the predictor's own footprint. The condition is

\[B_{\text{pred}} + (1-s)\,B_{\text{dense}} < B_{\text{dense}} \quad\Longleftrightarrow\quad B_{\text{pred}} < s\,B_{\text{dense}}\]

which reads: the predictor may cost up to the fraction it saves and not one byte more. For Llama-3-8B's feedforward stack at 11.3 GB in FP16 and 90 percent predicted sparsity that is about 10 GB of headroom, so a few hundred megabytes of low-rank predictor is comfortable and a predictor the size of a transformer layer is not. Spark Transformer refuses to add parameters at all, reallocating existing FFN parameters and attention key embeddings to serve as the predictor.

The second constraint is latency. Predict, gather, compute, serially, and the predictor joins the critical path at every one of 32 layers. Deja Vu's answer is to predict layer \(\ell+1\)'s sparsity from the input to layer \(\ell\), asynchronously, modelled on a hardware branch predictor. That is sound because a pre-norm block updates the residual stream additively, \(x_{\ell+1} = x_\ell + f_\ell(x_\ell)\), so once the accumulated norm dominates each block's update, adjacent layer inputs are nearly the same vector. It is weakest in early layers, where the stream is still being built.

Recall matters and precision does not

The predictor's two error types are wildly asymmetric. A false positive marks a dead neuron live: you read a row you did not need and produce a bit-identical output. A false negative drops a genuinely contributing neuron, which is noise at one in 14,336 and a quality regression if the bias is systematic, because a predictor that reliably misses the neurons firing on code or a low-resource language makes the model worse on exactly that traffic without moving aggregate perplexity.

Hence the design rule: tune for recall, over-predict the active set, and treat precision as a bandwidth optimisation rather than a correctness one.

Putting ReLU back, and specifying sparsity instead of hoping for it

Prediction and thresholding both work on the model you have. The alternative is to train for sparsity.

Relufication is the blunt version: replace SiLU or GELU with ReLU and continue pretraining. Every layer then shows sparsity above 90 percent, cutting inference computation by up to a factor of three with minimal quality loss (Mirzadeh et al., 2024, ReLU Strikes Back, ICLR 2024, arXiv:2310.04564). ProSparse softens the transition by ramping an \(L_1\) penalty on the intermediate activations along multi-stage sine curves and then shifting the activation threshold, reaching 89.32 percent on LLaMA2-7B and 87.89 percent on MiniCPM-1B at quality comparable to the Swish originals (Song et al., 2025, ProSparse, COLING 2025, arXiv:2402.13516).

The sharpest result in this line is a matched-sparsity comparison. TurboSparse applies ReLU to both branches of the gated block, a function the authors call dReLU, and at a forced 90 percent sparsity reports WikiText-2 perplexity of 29.19 where SwiGLU degrades to 112.36 (Song et al., 2024, Turbo Sparse, arXiv:2406.05955). Two easily conflated claims come apart there: a SwiGLU model can be sparsified, and it is catastrophically bad at it.

Q-Sparse goes further and makes the level a hyperparameter: a top-\(K\) mask keeps the \(K\) largest-magnitude activations in the forward pass, and because that mask has zero gradient almost everywhere, training pushes gradients through it with a straight-through estimator, the same device that makes quantisation-aware training work. The sparsity level is then known before training ends, so the engine can be built for a fixed shape rather than a measured distribution, and the approach holds for 1-bit models as well as full precision (Wang et al., 2024, Q-Sparse, arXiv:2407.10969).

[IMAGE: Four stacked density plots of FFN activation magnitudes on a log-x axis, for SwiGLU, relufied, ProSparse and dReLU models at equal parameter count, each with its chosen threshold as a vertical line. Caption: "SwiGLU has no mass at zero, only near it. Training for sparsity moves mass onto the spike."]

Seeing It in Motion

The predictor's placement in a decode step is the part most descriptions leave vague. It is a pipeline, and the point is that the prediction for the next layer overlaps the compute of the current one.

sequenceDiagram
    participant R as Runtime
    participant P as Predictor for layer 2
    participant K as Gather kernel
    participant G as GEMV
    R->>P: input to layer 1, async
    R->>K: layer 1 indices, ready earlier
    K->>G: 1434 gathered rows
    G->>R: layer 1 output
    P->>R: layer 2 index list
    Note over R,G: prediction for layer L plus 1 overlaps compute of layer L
    R->>K: layer 2 indices
    K->>G: 1434 gathered rows
    G->>R: layer 2 output
    Note over R,P: a false negative here is silent, a false positive only costs bytes

[IMAGE: Gantt-style trace of four decode layers, the predictor bar for layer L plus 1 overlapping the gather and GEMV bars for layer L, with the serial alternative drawn beneath. Caption: "Asynchronous prediction keeps the predictor off the critical path at all 32 layers."]

The second diagram is the decision that determines whether any of this is worth deploying.

flowchart TB
    S["Decode step"] --> Q1{"Batch size"}
    Q1 -- "one" --> Q2{"Do weights fit in fast memory"}
    Q1 -- "many" --> U["Union of active sets approaches dense"]
    U --> AH["Route attention heads instead of neurons"]
    AH --> W2["Up to 2.2x end to end"]
    Q2 -- "no" --> OFF["Offload: skipping avoids a slow transfer"]
    Q2 -- "yes" --> KER{"Neuron-aware kernels present"}
    OFF --> W1["4x to 25x versus naive loading"]
    KER -- "yes" --> W3["1.5x to 2x decode"]
    KER -- "no" --> N["No speedup at all"]

    classDef blue fill:#1e40af,stroke:#3b82f6,stroke-width:1px,color:#fff
    classDef purple fill:#6d28d9,stroke:#a78bfa,stroke-width:1px,color:#fff
    classDef emerald fill:#047857,stroke:#34d399,stroke-width:1px,color:#fff
    classDef rose fill:#be123c,stroke:#fb7185,stroke-width:1px,color:#fff
    classDef amber fill:#b45309,stroke:#fbbf24,stroke-width:1px,color:#fff
    class S blue
    class Q1,Q2,KER purple
    class W1,W2,W3 emerald
    class N rose
    class U,OFF,AH amber

Watch It Run

A decode step flowing left to right from hidden state through predictor to gather, small GEMV and scatter, with an animated self-loop on the predictor showing the look-ahead to the next layer and a feedback edge from the batch union back to the gather.
Solid animated edges carry the per-token data path: hidden state, predicted index list, gathered rows, small dense GEMV, scatter into the residual stream. The animated self-loop on the predictor is the look-ahead that predicts layer L plus 1 while layer L computes, and the feedback edge from the batch union is the path that erases the saving as batch size grows. The Mermaid figures above show the same structure if the animation is absent.

By the Numbers

Method Sparsity reported Quality cost Speed result Baseline
Deja Vu (2023) about 85 percent on OPT-175B; over 95 percent of MLP params per token flat to about 75 percent sparsity over 2x; over 6x FasterTransformer; Hugging Face
ReLU Strikes Back (2023) above 90 percent in all layers minimal, after continued pretraining up to 3x less computation the same model with SiLU or GELU
PowerInfer (2023) hot-cold power-law split accuracy retained up to 11.69x llama.cpp, one RTX 4090
LLM in a flash (2023) sparsity-aware loading plus windowing not the paper's axis 4x to 5x CPU, 20x to 25x GPU, at 2x DRAM size naive flash loading
CATS (2024) 50 percent within 1 to 2 percent of base, no fine-tuning about 15 percent faster generation dense, custom kernel
TEAL (2024) 40 to 50 percent model-wide minimal, Llama-⅔ and Mistral, 7B to 70B 1.53x at 40 percent, 1.8x at 50 percent dense decode, specialised kernels
TurboSparse (2024) close to 90 percent ppl 29.19 vs SwiGLU 112.36 at matched sparsity 2x to 5x; 11 tok/s for Mixtral-47B on a phone the dense SwiGLU model
Polar Sparsity (2025) attention heads, scalable under batching not compromised up to 2.2x end to end dense batched, Triton kernels
Spark Transformer (2025) 8 percent of FFN neurons, 256 attended tokens competitive on benchmarks 2.5x fewer FLOPs; 1.79x CPU, 1.40x GPU the dense Gemma-2 recipe

Sources, row by row: Liu et al., 2023; Mirzadeh et al., 2024; Song et al., 2024; Alizadeh et al., 2024; Lee et al., 2024; Liu et al., 2024; Song et al., 2024; Shrestha et al., 2025; Zhang et al., 2025. Every speed figure is the authors' own measurement against the baseline in the final column, on their hardware; the baselines differ enough that comparing speedups across rows is not meaningful, and each sparsity percentage depends on that paper's activation definition and quality budget.

A Concrete Example

Take Llama-3-8B on a single RTX 4090: 32 layers, \(d = 4096\), \(d_{\text{ff}} = 14336\), 32 query heads and 8 key-value heads, vocabulary 128,256. The card has 24 GB of GDDR6X on a 384-bit bus at 1,008 GB/s.

Step 1: where the bytes are. Each feedforward block holds three \(14336 \times 4096\) matrices, so \(3 \times 14336 \times 4096 = 176{,}160{,}768\) parameters, 352.3 MB in FP16, and 11.27 GB across 32 layers. Attention adds \(4096^2\) each for query and output and \(1024 \times 4096\) each for key and value: 41,943,040 parameters per layer, 2.68 GB in total. Embedding and output head are \(128256 \times 4096\) each, another 2.10 GB. Total 8.03 B parameters, 16.06 GB.

Step 2: the dense ceiling. A decode step at batch size one reads every weight once. At the card's peak 1,008 GB/s that is 15.9 ms, a bound of 63 tokens per second; at a realistic 85 percent of peak it is 18.7 ms and about 53 tokens per second.

Step 3: add 90 percent FFN activation sparsity with a predictor. Now each token needs 1,434 of the 14,336 neurons in each layer, so the FFN contribution falls from 11.27 GB to 1.13 GB. Budget a low-rank predictor per layer of \(4096 \times 512\) plus \(512 \times 14336\), which is 9.44 M parameters per layer, 302 M across the model, 0.60 GB in FP16. The bytes per step are now

\[1.13 + 0.60 + 2.68 + 2.10 = 6.51 \text{ GB}\]

against 16.06 GB dense, a 2.47x reduction in traffic. The predictor consumed 6 percent of the budget it unlocked, well inside the inequality from earlier.

Step 4: subtract the kernel tax. Gathered rows do not stream like contiguous ones. Hold the dense portions at 85 percent of peak and let the gathered FFN reads achieve 55 percent: \(5.38/857 + 1.13/554 = 6.3 + 2.0 = 8.3\) ms, about 120 tokens per second. That is 2.25x over the realistic dense baseline, consistent with TurboSparse's 2x to 5x range at similar sparsity, and already a full factor below the naive intuition.

Step 5: now serve 32 users. Each token in the batch picks its own 1,434 neurons and the layer must read the union. If selections were independent, each neuron escapes all 32 tokens with probability \(0.9^{32} \approx 0.034\), so the union covers 96.6 percent of the layer. Selections are not independent, which helps. Suppose 1,000 of each token's active neurons come from a hot pool of 2,000 and the remaining 434 land anywhere in the other 12,336:

\[2{,}000 + 12{,}336\Big(1 - \big(1 - \tfrac{434}{12{,}336}\big)^{32}\Big) \approx 2{,}000 + 8{,}416 = 10{,}416\]

which is 73 percent of the layer. Per-token sparsity of 90 percent has become batch sparsity of 27 percent. In bytes per generated token, the dense batched path reads \(16.06/32 = 0.50\) GB; the sparse path reads \((8.23 + 0.60 + 2.68 + 2.10)/32 = 0.43\) GB. A 15 percent byte saving, paid for with gather overhead on every layer and with the loss of a clean dense GEMM that a batch of 32 would have fed efficiently.

That is the argument in five steps: worth 2.25x to one user and roughly nothing to thirty-two.

[IMAGE: Line chart, batch size 1 to 64 on a log x-axis, with effective batch sparsity falling from 90 percent to under 20 percent and speedup over dense falling from 2.2x through 1.0x, the crossover marked. Caption: "The crossover, not the peak, decides a deployment."]

Where It Breaks

Decode is bandwidth-bound, so FLOP reductions are the wrong currency

At batch size one a decode step performs roughly one multiply-accumulate per weight read. That arithmetic intensity is orders of magnitude below what a modern accelerator needs to saturate its arithmetic units, so elapsed time tracks bytes moved and not operations performed (the mechanics are in arithmetic intensity and memory-bound deep learning). A method that halves FLOPs without halving bytes has bought nothing. This is why the saving must be realised at the gather, before the read, and why "2.5x fewer FLOPs" and "1.40x faster" are both true of the same system.

The gather is not free, and the down projection is worse than the up

Reading 1,434 scattered rows of a row-major up projection is tolerably coalesced; each row is 8 KB of contiguous FP16. The matching 1,434 columns of the down projection are strided, touching one element per row of a 14,336-wide matrix, which wastes most of each cache line unless the weights are stored transposed for the purpose. Add index traffic and per-layer kernel launches and gathered reads land well below the dense path's efficiency. Published wall-clock numbers include this tax; sparsity percentages do not.

Peak throughput also comes from fixed-shape tiles with regular access, and an arbitrary 1,434-row subset does not tile, the same reason N:M semi-structured sparsity exists. Whole rows are better than scattered zeros, and still irregular in a way the hardware dislikes.

Batching erases the premise

Serving systems batch aggressively for exactly the reason sparsity exists: to raise arithmetic intensity and escape the bandwidth bound. Those are the same goal pursued by incompatible means, so a throughput-oriented deployment cannot have both. Polar Sparsity names the resolution: as batch size and sequence length grow, MLP layers become more compute-efficient under batching while their sparsity vanishes, and the exploitable selectivity migrates to attention heads, which stay selective because each request attends to its own context. Routing heads with learned routers and Triton kernels gives up to 2.2x end to end on OPT and Llama-⅔.

Prefill gets nothing

Prompt processing runs hundreds or thousands of positions through the same weights at once, which is the union problem at a much larger batch, and it is compute-bound anyway. Activation sparsity is a decode optimisation, so for a retrieval-augmented request with a long prompt and a short answer the share of latency it can touch may be small.

Perplexity hides the damage, and the calibration expires

Thresholds and predictors are calibrated against a quality budget, and perplexity is the cheapest budget to hold. A perplexity-neutral configuration can still lose several points on multi-step reasoning, the same asymmetry that shows up in depth pruning: recalling a fact tolerates perturbation, chaining many steps feeds each error forward. Evaluating a sparsified model therefore means free-form generation and multi-step tasks, not a benchmark average.

Worse, the calibration is a per-checkpoint artefact. Fine-tune, merge an adapter or continue post-training and the activation statistics move. Refitting is cheap, but it has to be an automated release step or the sparsity decays into systematic false negatives that nothing in CI is watching for.

Alternative Designs

Design How it works Key advantage Key limitation Best when
Activation sparsity Skip rows and columns per token, by predictor or threshold No quality loss in principle, no retraining in the free variants Dies under batching, needs custom kernels, decode only Batch size one, especially with offloading
Weight-only quantisation Store weights at 4 bits, dequantise in the kernel Unconditional byte reduction that survives batching Quality cost at low bit widths, no FLOP reduction Almost always the first thing to try
Mixture of Experts Train the routing, activate a few contiguous experts per token Conditional compute with kernel-friendly structure Total memory stays large, and it is not a retrofit Choosing an architecture, not optimising a checkpoint
Depth pruning Delete whole transformer blocks permanently Near-linear batch-one speedup, no special kernels Permanent quality loss, concentrated in reasoning A lower fixed quality bar is acceptable
N:M weight sparsity Zero two of every four weights in a fixed pattern Hardware support on recent NVIDIA parts Retraining needed, realised speedups well below 2x The hardware supports it and training budget exists
Speculative decoding Draft several tokens cheaply, verify in one pass Multiplies tokens per weight read, and survives batching Needs a draft model, gains track acceptance rate Latency-sensitive serving at moderate batch

Two deserve direct comparison. Quantisation spends the same budget, bytes moved per decode step, so the two are partly substitutes whose gains do not multiply, and a 4-bit dense model is often the simpler route to the same reduction with no gather kernels and no predictor to maintain (weight-only post-training quantisation). Mixture of Experts is activation sparsity with the routing decision made explicit, trained, and arranged so the active set is contiguous (mixture of experts): the same win, taken structurally, with kernel support behind it.

[IMAGE: Stacked bars of bytes read per decode step for Llama-3-8B under FP16 dense, FP16 plus 90 percent sparsity, 4-bit dense, 4-bit plus sparsity and an MoE equivalent, segmented by FFN, attention, embeddings and predictor. Caption: "Sparsity and quantisation compete for the same segment."]

How It Is Used in Practice

The deployed successes share a shape: one user at a time, and a cliff in the memory hierarchy.

Apple's work is the clearest case. Storing parameters in flash and paging them into DRAM on demand makes every avoided neuron an avoided transfer across a slow link rather than a cheap read, and the system runs models up to twice the size of available DRAM, 4x to 5x faster than naive loading on CPU and 20x to 25x on GPU. Two techniques do the work: windowing, which reuses neurons activated by recent tokens instead of reloading them, and row-column bundling, which stores each neuron's up-projection row next to its down-projection column so one sequential read fetches both (Alizadeh et al., 2024, LLM in a flash, ACL 2024, arXiv:2312.11514). The second is the practical answer to the strided-column problem above, and it is a storage-layout decision rather than a kernel one.

PowerInfer is the consumer-GPU case, and its 11.69x over llama.cpp on one RTX 4090 is a statement about offloading policy rather than sparse kernels: when a model does not fit in 24 GB the question is which parameters live where, and activation frequency is an excellent answer. TurboSparse pushes the same logic onto phones, with Mixtral-47B at around 11 tokens per second on a mobile device, where there is no batching to lose.

Datacentre serving has not adopted MLP activation sparsity, and the worked example explains why. What has crossed over is the attention-head form, where selectivity survives batching, and the predictor-as-router idea, now recognisable as a less-structured sibling of MoE routing. Spark Transformer shows where the technique belongs in a modern recipe: 8 percent of FFN neurons active, at most 256 tokens attended per query, parameters reallocated rather than added.

[IMAGE: Three-panel schematic of the hierarchies that make sparsity pay, bandwidth drawn to scale: flash to DRAM, CPU DRAM to GPU over PCIe, GPU HBM, with the avoided transfer highlighted in each. Caption: "The steeper the cliff below the resident set, the more a skipped neuron is worth."]

Insights Worth Remembering

  1. Activation sparsity does not make a model smaller, it makes a token cheaper. Every weight is still required, because some other token needs it. Storage is unchanged and only per-token read cost falls, the exact inverse of pruning, which is why the two compose rather than compete.

  2. The ordering problem is the field. Knowing a neuron is dead requires computing it, so every method is a way of finding out early: predict it, threshold the layer's input, or decide statically from activation frequency. Choosing among those three is choosing a system.

  3. Nominal sparsity and elapsed time have a poor exchange rate. Fifty percent sparsity bought CATS about 15 percent latency; a 2.5x FLOP cut bought Spark Transformer 1.40x on GPU. A claim phrased purely as a sparsity percentage has not told you the thing you need.

  4. Batching and sparsity are competing answers to the same problem. Both exist to escape the bandwidth bound of batch-one decode, so succeeding at one forecloses the other. Hence throughput-oriented serving skipped this technique, and the sparsity that survives batching lives in attention heads.

  5. Recall, not accuracy, is the predictor metric that matters. A false positive costs bytes; a false negative costs output quality, silently and selectively. Over-predicting is the correct bias, and a predictor judged on aggregate perplexity has not been judged.

  6. The activation function was an inference decision all along. Swapping ReLU for SiLU bought small quality gains and sold a 90 percent per-token bandwidth saving, a trade TurboSparse prices retrospectively at 29.19 against 112.36 perplexity. It also means sparsity is only as real as its activation definition, since on a SwiGLU model the exact-zero count is nil.

  7. A sparse checkpoint on a dense runtime is just a model. The engine is the deliverable: gather kernels, index management, a storage layout that puts each neuron's row and column together. Absent those, 89 percent sparsity is worth zero percent speedup.

Open Questions

Does activation sparsity grow with scale fast enough to matter at frontier sizes? Measured: sparsity strengthens with width and depth, and a 2025 cross-family sweep finds tolerable sparsity rising with model size while surviving instruction tuning and reasoning post-training (Haziza et al., 2025, arXiv:2509.00454). Not established: whether that holds for the largest MoE models, where each expert is already a narrow FFN and the conditional computation has been spent once.

Can the union problem be designed around rather than conceded? Polar Sparsity changes substrate, routing attention heads instead of neurons. An alternative is to batch tokens by predicted active set so each micro-batch shares neurons, which is scheduling rather than kernel work. No published system does this at scale.

Is the right final form of this idea simply MoE? An MoE layer is trained conditional computation with contiguous active sets and mature kernels, and the argument for activation sparsity as a separate technique rests on it applying to checkpoints that already exist. Whether anyone will train a new dense model expecting to sparsify it is a question about the next generation of recipes, not about the method.

How much of the published quality-neutrality survives harder evaluation? The literature leans on perplexity and multiple-choice benchmarks. Whether 50 percent training-free sparsity is genuinely free on agentic tool use, long-horizon code editing and multi-step reasoning is, on published evidence, untested.

Will hardware meet this halfway? N:M weight sparsity exists on current accelerators because the pattern was constrained enough to build for. A fixed block or N:M structure on the active set would tile and could reach a far better conversion rate. Several 2025 papers propose variants; none has silicon behind it.

Sources and Further Reading

The phenomenon

  1. Li, Z., et al. (2022). "The Lazy Neuron Phenomenon: On Emergence of Activation Sparsity in Transformers." arXiv:2210.06313
  2. Luo, Y., Song, C., Han, X., et al. (2024). "Sparsing Law: Towards Large Language Models with Greater Activation Sparsity." arXiv:2411.02335
  3. Haziza, D., et al. (2025). "Universal Properties of Activation Sparsity in Modern Large Language Models." arXiv:2509.00454

Exploiting it at inference

  1. Liu, Z., Wang, J., Dao, T., et al. (2023). "Deja Vu: Contextual Sparsity for Efficient LLMs at Inference Time." ICML 2023, PMLR 202:21631-21657. arXiv:2310.17157
  2. Lee, J., et al. (2024). "CATS: Contextually-Aware Thresholding for Sparsity in Large Language Models." arXiv:2404.08763
  3. Liu, J., Ponnusamy, P., Cai, T., et al. (2024). "Training-Free Activation Sparsity in Large Language Models." ICLR 2025, spotlight. arXiv:2408.14690
  4. Shrestha, S., et al. (2025). "Polar Sparsity: High Throughput Batched LLM Inferencing with Scalable Contextual Sparsity." arXiv:2505.14884

Training for it

  1. Mirzadeh, I., Alizadeh, K., et al. (2024). "ReLU Strikes Back: Exploiting Activation Sparsity in Large Language Models." ICLR 2024. arXiv:2310.04564
  2. Song, C., Han, X., Zhang, Z., et al. (2025). "ProSparse: Introducing and Enhancing Intrinsic Activation Sparsity within Large Language Models." COLING 2025. arXiv:2402.13516
  3. Song, Y., Xie, H., Zhang, Z., et al. (2024). "Turbo Sparse: Achieving LLM SOTA Performance with Minimal Activated Parameters." arXiv:2406.05955
  4. Wang, H., Ma, S., Wang, R., & Wei, F. (2024). "Q-Sparse: All Large Language Models can be Fully Sparsely-Activated." arXiv:2407.10969
  5. Zhang, Z., et al. (2025). "Spark Transformer: Reactivating Sparsity in FFN and Attention." arXiv:2506.06644

Systems and offloading

  1. Song, Y., Mi, Z., Xie, H., & Chen, H. (2024). "PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU." SOSP 2024. arXiv:2312.12456
  2. Alizadeh, K., Mirzadeh, I., Belenko, D., et al. (2024). "LLM in a flash: Efficient Large Language Model Inference with Limited Memory." ACL 2024. arXiv:2312.11514

Free to read, no ads, no sign-up. If it was useful you can buy me a coffee.