RL for Language Models
RLHF as an RL problem, KL-regularised objectives, GRPO, RLVR, and reward over-optimisation.
—
0 known
0 to review
Loading deck…
Space to flipSpace flip · ← → move · K known · R review · S shuffle