Training & Alignment
21 min
RL from Verifiable Rewards: Training Models on Answers That Can Be Checked
Replace the reward model with a function that simply checks the answer, and a frontier reasoning model falls out of pure reinforcement learning. The catch is what 'checkable' quietly assumes, and what the model learns to exploit.