RL for Language Models

RLHF as an RL problem, KL-regularised objectives, GRPO, RLVR, and reward over-optimisation.

0 known 0 to review

Loading deck…

Space to flip

Space flip · move · K known · R review · S shuffle