Flashcards
11,105 cards in 101 decks, written from the same source as the concepts. Space to flip, arrows to move, K and R to sort what you know from what you do not. Progress lives in your browser.
Efficiency, Compression & Edge AI
Making a model smaller, cheaper and local without giving away the thing that made it useful.
Quantisation
Post-training and quantisation-aware methods, outlier channels, GPTQ and AWQ, and low-bit arithmetic formats.
Knowledge Distillation
Soft targets and temperature, sequence-level and on-policy distillation, and when a student beats its teacher.
Sparsity & Pruning
Magnitude and second-order criteria, structured versus unstructured sparsity, and the hardware that rewards it.
Efficient Architectures
Small language models, depth-width tradeoffs, weight sharing, and architectures designed for a latency budget.
On-Device & Edge AI
Mobile NPUs, memory-bound inference on consumer silicon, compilation targets and privacy-driven local models.