Inference & Serving
27 min
Activation Sparsity: The 90 Percent of a Dense Model That Does Nothing Per Token
Feed one token through T5-Base and 97 percent of its MLP neurons output exactly zero. Nobody pruned the model; the sparsity arrived on its own. The honest exchange rate for cashing it in is a 2.5x cut in arithmetic for 1.40x on a…