Sparsity & Pruning advanced 8 min read 7 flashcards

Cashing In Activation Sparsity on Real Hardware

Fifty percent activation sparsity converts to roughly fifteen percent latency and a 2.5x FLOP cut converts to 1.40x on a GPU, because decode is bandwidth-bound, gathers are not free, and batching makes the union of active neurons dense.

Two published pairs of numbers explain this entire topic. CATS reports 50 percent activation sparsity and about 15 percent improvement in token-generation latency (Lee et al., 2024, arXiv:2404.08763). Spark Transformer reports a 2.5x reduction in FLOPs and up to 1.40x decoding speedup on GPU (Zhang et al., 2025, arXiv:2506.06644). Neither is a disappointing result; both are among the better ones in the literature. The gap between nominal sparsity and elapsed time is the subject.

Why the conversion rate is poor

Decode is bandwidth-bound, so only bytes count. At batch size one, each generated token multiplies the weights by a single vector. Arithmetic intensity is roughly one multiply-accumulate per weight read, orders of magnitude below the ratio a modern accelerator needs to saturate its arithmetic units, so time is set by weight traffic and not by FLOPs (see arithmetic intensity and memory-bound deep learning). A FLOP reduction that does not reduce bytes moved is worth nothing here, which is why sparsity has to be exploited at the gather, before the read, rather than by skipping multiplications.

Gathers cost real time. A dense GEMV streams contiguous memory at near peak. A sparse one reads an index list, gathers scattered rows, and runs a smaller matrix multiply on a non-standard shape. Row gathers from a row-major up projection are reasonably coalesced; the matching column gather from the down projection is strided and much less so. Both pay index traffic and kernel-launch overhead that the dense path does not.

Tensor cores want dense tiles. Peak throughput comes from fixed-shape tiles fed by regular access patterns. Arbitrary subsets of rows do not tile, which is the same reason unstructured weight sparsity rarely beats dense matmul and why N:M semi-structured sparsity exists.

Prefill is already dense. Prompt processing is compute-bound and batched over positions. The union of active neurons across a thousand positions is essentially everything, so sparsity applies to decode only, and the end-to-end benefit is diluted by whatever share of the request prefill occupies.

The union problem, quantified

This is the one that ends most deployments. Each token in a batch selects its own active set, and the layer must read the union of all of them.

Take a 14,336-wide feedforward layer at 90 percent per-token sparsity, so 1,434 neurons per token. If selections were independent and uniform, the expected union over a batch of 32 would be \(1 - 0.9^{32} \approx 96.6\) percent of the layer. Real selections are not independent, because activation follows the power law PowerInfer exploits: suppose 1,000 of each token's active neurons come from a hot pool of 2,000 and the other 434 land anywhere in the remaining 12,336. The union is then \(2{,}000 + 12{,}336\big(1 - (1 - 434/12{,}336)^{32}\big) \approx 10{,}400\) neurons, about 73 percent of the layer. Per-token sparsity of 90 percent has become batch sparsity of 27 percent, and you are paying gather overhead to achieve it.

Polar Sparsity makes this the central finding rather than a caveat: as batch size and sequence length grow, MLP layers become more compute-efficient under batching while their sparsity vanishes, and the exploitable sparsity migrates to attention heads, which stay selective because each request attends to its own context. Routing heads rather than neurons, with custom Triton kernels, delivers up to 2.2x end-to-end speedups on OPT and Llama-⅔ across batch sizes (Shrestha et al., 2025, arXiv:2505.14884).

Where it genuinely pays

The clean wins share a shape: batch size one, and a memory hierarchy with a cliff in it.

Offloading. When weights do not fit in fast memory, skipping a neuron avoids a transfer across a slow link rather than a cheap read. Apple's flash-resident inference runs models up to twice the size of available DRAM, with 4x to 5x faster inference than naive loading on CPU and 20x to 25x on GPU, by combining sparsity awareness with windowing that reuses recently activated neurons and row-column bundling that turns scattered reads into large sequential ones (Alizadeh et al., 2024, LLM in a flash, ACL 2024, arXiv:2312.11514).

Consumer GPUs with a CPU alongside. PowerInfer's hot-cold partition keeps frequently activated neurons resident on the GPU and runs the rest on the CPU, outperforming llama.cpp by up to 11.69x on a single RTX 4090 (Song et al., 2024, SOSP 2024, arXiv:2312.12456).

Phones and CPUs. TurboSparse-Mixtral-47B at around 11 tokens per second on a mobile device is a workload with no batching to lose (Song et al., 2024, arXiv:2406.05955).

Datacentre serving is the opposite case. Throughput-oriented deployments batch aggressively precisely because batching converts a bandwidth-bound problem into a compute-bound one, and that conversion is what destroys the sparsity.

When it breaks

A speedup is only as meaningful as its baseline. An 11.69x figure against llama.cpp on a 24 GB card is a statement about offloading, not about sparse kernels. Compare against the strongest dense engine that fits the same hardware.

The predictor is on the critical path unless you hide it. Synchronous prediction adds latency per layer; asynchronous look-ahead removes it at the cost of an approximation that is weakest in early layers.

Sparsity and quantisation compete for the same budget. Both reduce bytes per token. Applying one shrinks the headroom available to the other, so gains do not compose multiplicatively, and a 4-bit dense model is often the simpler way to get the same bandwidth reduction (see weight-only post-training quantisation).

MoE already took this win, structurally. A mixture-of-experts layer is activation sparsity with the routing decision made explicit, trained, and arranged so that the active set is contiguous and kernel-friendly (see mixture of experts). For a team choosing an architecture rather than optimising a fixed checkpoint, that is the better-supported path.

References and further reading

Every source this page cites, in the order it cites them. All of them open in a new tab.

  1. Lee et al., 2024, arXiv:2404.08763 arxiv.org
  2. Zhang et al., 2025, arXiv:2506.06644 arxiv.org
  3. Shrestha et al., 2025, arXiv:2505.14884 arxiv.org
  4. Alizadeh et al., 2024, LLM in a flash, ACL 2024, arXiv:2312.11514 arxiv.org
  5. Song et al., 2024, SOSP 2024, arXiv:2312.12456 arxiv.org
  6. Song et al., 2024, arXiv:2406.05955 arxiv.org
Check yourself

7 flashcards for this concept

Click a card to reveal the answer.

Drill the whole track