Flashcards
11,105 cards in 101 decks, written from the same source as the concepts. Space to flip, arrows to move, K and R to sort what you know from what you do not. Progress lives in your browser.
All decks
01Foundations
02Transformer Internals
03Training & Fine-Tuning
04Reinforcement Learning
05Inference, Systems & Hardware
06Applied LLM Engineering
07Reasoning, Evaluation & Safety
08Multimodal & Applications
09Classical ML & Statistical Learning
10Causal Inference & Experimentation
11Time Series & Forecasting
12Graphs, Recommenders & Structured Data
13Generative Modelling Beyond Transformers
14Efficiency, Compression & Edge AI
15Search & Information Retrieval
16Data & Feature Engineering
17MLOps & Platform Engineering
18Security, Privacy & Adversarial ML
19Governance, Risk & Responsible AI
20Human-AI Interaction, Product & Economics
04
Reinforcement Learning
Classical RL, then the specific dialect of it that post-trains language models.
5decks
489cards
RL Foundations
MDPs, value functions, TD learning, policy gradients, actor-critic, TRPO and PPO.
RL for Language Models
RLHF as an RL problem, KL-regularised objectives, GRPO, RLVR, and reward over-optimisation.
Bandits & Exploration
Regret, UCB, Thompson sampling, contextual bandits, and exploration under a budget.
Offline & Model-Based RL
Learning from logged data, distribution shift, conservative value estimation, world models and planning.
Multi-Agent RL
Self-play, equilibria, credit assignment across agents, emergent coordination and non-stationarity.