Flashcards
2,671 cards in 36 decks, written from the same source as the concepts. Space to flip, arrows to move, K and R to sort what you know from what you do not. Progress lives in your browser.
07
Reasoning, Evaluation & Safety
Models that think longer, the evals that measure them, and the failure modes that matter.
3decks
157cards
Reasoning Models
Test-time compute, process reward models, the o-series, DeepSeek-R1 and contamination.
Evaluation & MLOps
Benchmarks, LLM-as-judge, red-teaming, model registries, drift detection and observability.
Safety & Alignment
Prompt injection, jailbreaks, Constitutional AI, reward hacking and mechanistic interpretability.