RL for Language Models
RLHF as an RL problem, KL-regularised objectives, GRPO, RLVR, and reward over-optimisation.
21concepts
142flashcards
172minutes of reading
No beginner concepts in this track. Show all.
RLHF as an RL problem, KL-regularised objectives, GRPO, RLVR, and reward over-optimisation.
No beginner concepts in this track. Show all.