RL for Language Models

RLHF as an RL problem, KL-regularised objectives, GRPO, RLVR, and reward over-optimisation.

21concepts
142flashcards
172minutes of reading

No beginner concepts in this track. Show all.