Relufication and Sparsity-Aware Pretraining
Swapping SiLU back to ReLU and continuing to pretrain recovers 90 percent activation sparsity that SwiGLU models never had, and the explicit top-K and dReLU variants show the quality cost of that sparsity is close to zero if you train for it.
Around 2020 the field traded ReLU for GELU and SiLU on the strength of small, consistent quality gains. The bill arrived at inference time: the new activations have no exact zeros, so an inference engine cannot skip anything, and the 90-plus percent sparsity that ReLU models hand over for free was gone. Apple's argument is that the trade was worth revisiting, because the sparsity is worth more than the quality delta it bought (Mirzadeh et al., 2024, ReLU Strikes Back: Exploiting Activation Sparsity in Large Language Models, ICLR 2024, arXiv:2310.04564).
The procedure they call relufication is blunt: replace the activation function with ReLU, then continue pretraining. All layers then show sparsity above 90 percent, cutting inference computation by up to a factor of three with minimal quality loss. The cost is the continued-pretraining run, measured in hundreds of billions of tokens.
Getting there without wrecking the distribution
A naive swap moves every activation distribution in the network at once, and quality dips before the continued pretraining recovers it. ProSparse makes the transition gradual: substitute ReLU, then apply an \(L_1\) regularisation on the intermediate activations whose coefficient ramps up along a multi-stage sine curve, then shift the activation threshold. That reaches 89.32 percent sparsity on LLaMA2-7B, 88.80 percent on LLaMA2-13B and 87.89 percent on MiniCPM-1B, with performance comparable to the original Swish-activated models (Song et al., 2025, ProSparse: Introducing and Enhancing Intrinsic Activation Sparsity within Large Language Models, COLING 2025, arXiv:2402.13516).
A different line keeps the gate structure and changes the non-linearity. TurboSparse's dReLU applies a ReLU to both branches of the gated block, and models trained with it reach close to 90 percent sparsity while matching SwiGLU quality. The cleanest number in that paper is a controlled comparison at matched sparsity: forced to 90 percent, dReLU gives WikiText-2 perplexity 29.19 where SwiGLU collapses to 112.36 (Song et al., 2024, Turbo Sparse: Achieving LLM SOTA Performance with Minimal Activated Parameters, arXiv:2406.05955). Sparsified Mistral-7B activates 2.5B parameters per iteration and Mixtral-47B activates 4.3B, with 2x to 5x decoding speedups and TurboSparse-Mixtral-47B running at about 11 tokens per second on a phone.
Making sparsity a trained constraint
Relufication induces sparsity and hopes for a useful level. The alternative is to specify the level. Q-Sparse applies top-\(K\) sparsification to activations in the forward pass and uses a straight-through estimator for the backward pass, exactly the trick that makes quantisation-aware training differentiable through a step function:
where \(M_K\) keeps the \(K\) largest-magnitude entries. Because \(K\) is a hyperparameter rather than an emergent property, the sparsity level is known before training finishes and the engine can be built for it. Q-Sparse works from scratch, as continued training of an existing model, and as fine-tuning, holds for 1-bit models as well as full precision, and comes with an inference-optimal scaling law relating model size to sparsity (Wang et al., 2024, Q-Sparse: All Large Language Models can be Fully Sparsely-Activated, arXiv:2407.10969).
Spark Transformer takes the same specify-it approach into a modern recipe, using a statistical top-\(k\) that accelerators can execute and reallocating existing FFN parameters and attention key embeddings to serve as the low-cost predictor rather than adding a separate network. Pretrained with the Gemma-2 recipe it activates 8 percent of FFN neurons and attends to at most 256 tokens per query, for a 2.5x FLOP reduction and decoding speedups of up to 1.79x on CPU and 1.40x on GPU (Zhang et al., 2025, Spark Transformer: Reactivating Sparsity in FFN and Attention, arXiv:2506.06644).
That last pair of numbers is the honest summary of the field: a 2.5x reduction in arithmetic yields 1.40x on a GPU.
When it breaks
The training bill is the whole argument. Relufication and ProSparse both need continued pretraining on a large token budget. If you are not already running such a job, this is not a cheap optimisation, and training-free sparsification at 40 to 50 percent may be the better trade.
Sparsity targets interact with quality unevenly across tasks. The published comparisons lean on perplexity and multiple-choice benchmarks. Free-form generation and multi-step reasoning are the places where induced sparsity is most likely to hurt and least likely to be measured.
An inference engine is part of the deliverable. A ProSparse checkpoint on a stock dense runtime is just a slightly different model with no speed advantage at all. The sparsity only pays through neuron-aware kernels, and the paper's acceleration results depend on them.
Specified sparsity is a commitment. Choosing \(K\) fixes the engine's shape. Raising sparsity later is another training run, not a configuration change.
References and further reading
Every source this page cites, in the order it cites them. All of them open in a new tab.
- Mirzadeh et al., 2024, ReLU Strikes Back: Exploiting Activation Sparsity in Large Language Models, ICLR 2024, arXiv:2310.04564 arxiv.org
- Song et al., 2025, ProSparse: Introducing and Enhancing Intrinsic Activation Sparsity within Large Language Models, COLING 2025, arXiv:2402.13516 arxiv.org
- Song et al., 2024, Turbo Sparse: Achieving LLM SOTA Performance with Minimal Activated Parameters, arXiv:2406.05955 arxiv.org
- Wang et al., 2024, Q-Sparse: All Large Language Models can be Fully Sparsely-Activated, arXiv:2407.10969 arxiv.org
- Zhang et al., 2025, Spark Transformer: Reactivating Sparsity in FFN and Attention, arXiv:2506.06644 arxiv.org
6 flashcards for this concept
Click a card to reveal the answer.