Flashcards
11,105 cards in 101 decks, written from the same source as the concepts. Space to flip, arrows to move, K and R to sort what you know from what you do not. Progress lives in your browser.
All decks
01Foundations
02Transformer Internals
03Training & Fine-Tuning
04Reinforcement Learning
05Inference, Systems & Hardware
06Applied LLM Engineering
07Reasoning, Evaluation & Safety
08Multimodal & Applications
09Classical ML & Statistical Learning
10Causal Inference & Experimentation
11Time Series & Forecasting
12Graphs, Recommenders & Structured Data
13Generative Modelling Beyond Transformers
14Efficiency, Compression & Edge AI
15Search & Information Retrieval
16Data & Feature Engineering
17MLOps & Platform Engineering
18Security, Privacy & Adversarial ML
19Governance, Risk & Responsible AI
20Human-AI Interaction, Product & Economics
05
Inference, Systems & Hardware
Where the model meets the silicon, the memory bus and the latency budget.
4decks
645cards
Inference Optimisation
KV cache, FlashAttention, speculative decoding, quantisation and continuous batching.
Accelerator Architecture
The memory wall, roofline analysis, GPU execution model, interconnects and systolic arrays.
Kernels & Compilers
CUDA, Triton, fusion, tiling, torch.compile, CUDA graphs and roofline-guided optimisation.
Serving Systems
Prompt caching, gateways and routing, token accounting, and multi-tenant isolation.