14

Efficiency, Compression & Edge AI

Making a model smaller, cheaper and local without giving away the thing that made it useful.

5tracks
25concepts
308cards
3.0hreading
Quantisation Post-training and quantisation-aware methods, outlier channels, GPTQ and AWQ, and low-bit arithmetic formats. 5 concepts · 64 cards
Knowledge Distillation Soft targets and temperature, sequence-level and on-policy distillation, and when a student beats its teacher. 5 concepts · 64 cards
Sparsity & Pruning Magnitude and second-order criteria, structured versus unstructured sparsity, and the hardware that rewards it. 5 concepts · 62 cards
Efficient Architectures Small language models, depth-width tradeoffs, weight sharing, and architectures designed for a latency budget. 5 concepts · 58 cards
On-Device & Edge AI Mobile NPUs, memory-bound inference on consumer silicon, compilation targets and privacy-driven local models. 5 concepts · 60 cards