Training-Free Activation Sparsification
A magnitude threshold calibrated per layer gets 40 to 50 percent model-wide activation sparsity out of an off-the-shelf SwiGLU model with no retraining and no predictor network, which is a smaller win than relufication and costs nothing to try.
Relufication and sparsity-regularised pretraining both reach roughly 90 percent activation sparsity, and both cost hundreds of billions of tokens of continued training. For a team that holds weights rather than a training cluster, the relevant question is what sparsity is available for free.
The answer, as of early 2026, is 40 to 50 percent model-wide, with no training at all and no auxiliary predictor. TEAL applies a plain magnitude threshold to hidden states throughout the model and reports 40 to 50 percent sparsity with minimal degradation across Llama-2, Llama-3 and Mistral from 7B to 70B, turning that into wall-clock decoding speedups of up to 1.53x at 40 percent and 1.8x at 50 percent using specialised kernels (Liu et al., 2024, Training-Free Activation Sparsity in Large Language Models, arXiv:2408.14690, ICLR 2025 spotlight).
Why a threshold works at all
The justification is distributional. In Llama-architecture models the entries of the hidden states entering each projection are close to zero-mean and unimodal. For such a distribution, zeroing every entry whose magnitude is below \(\tau\) removes a contiguous band around the mode, which is where the entries that contribute least live. Write the sparsified input as
and the induced error in the layer output is \(W(x - \tilde{x})\), a sum of many terms each individually small. The threshold is set per layer, not globally, because the scale of the hidden states differs by depth and a single \(\tau\) would over-sparsify some layers while barely touching others.
CATS makes the same move one step earlier in the block, thresholding the output of the gate projection in a SwiGLU FFN using a cutoff read off a calibration quantile, and reports downstream performance within about one to two percent of the base model at 50 percent activation sparsity with no fine-tuning (Lee et al., 2024, CATS: Contextually-Aware Thresholding for Sparsity in Large Language Models, arXiv:2404.08763).
Threshold against predictor
A threshold has no training cost, no extra parameters, and no distribution-shift exposure beyond the calibration set used to pick \(\tau_\ell\). Its weakness is the ordering problem from contextual sparsity: CATS must compute the full gate projection to evaluate its own threshold, so the saving applies to the up and down projections and not to the gate. That caps the achievable fraction at roughly two thirds of the block even at perfect sparsity. TEAL avoids some of this by thresholding the input hidden state rather than an intermediate, which sparsifies the columns of every projection it feeds, including the gate.
The published wall-clock numbers show how much is lost between the two framings. CATS at 50 percent activation sparsity reports about 15 percent improvement in token-generation latency with its custom kernel. TEAL at the same nominal sparsity reports 1.8x. Same headline sparsity, very different results, and the difference is entirely which matrices the sparsity actually applies to and how well the kernel exploits it.
Calibrating the thresholds
Three decisions matter more than the threshold rule itself.
The quality objective. Choosing \(\tau_\ell\) to hold perplexity within a budget and choosing it to hold a downstream benchmark are different optimisations, and perplexity is the more forgiving. A perplexity-neutral threshold can still cost several points on multi-step reasoning, the same asymmetry that shows up in depth pruning.
The allocation across layers. A uniform sparsity target is simple and leaves value on the table, because tolerance varies by depth. Greedy per-layer search against a held-out quality metric recovers most of it, in direct analogy to mixed-precision bit allocation.
The calibration corpus. A few hundred sequences are enough to estimate a quantile stably, but they must look like production traffic. Thresholds calibrated on English prose and deployed on a code assistant are a reasonable way to ship a silent regression.
When it breaks
You cannot threshold what you have not computed. Any method that reads a projection's output to decide whether to read the projection has already spent the bandwidth. Sparsifying the input side is strictly more useful than sparsifying an intermediate.
50 percent sparsity is not a 2x speedup. The relationship between zeroed activations and elapsed time runs through gather overhead, kernel efficiency and memory-boundedness, and the conversion rate is poor. Expect well under half the nominal saving without a dedicated kernel.
Compatibility with quantisation is not automatic but has been demonstrated. TEAL reports composing with weight quantisation for further gains; the two interact through the same bandwidth budget, so they are partly substitutes, and stacking them does not multiply cleanly.
Thresholds drift with fine-tuning. Any post-training that moves activation scales invalidates the calibration. Recalibration is cheap, but it has to actually happen in the pipeline.
References and further reading
Every source this page cites, in the order it cites them. All of them open in a new tab.
6 flashcards for this concept
Click a card to reveal the answer.