Training & Alignment
25 min
Why Policy Gradients Need a Baseline: Variance, Trust Regions, and the Road to PPO
Add 1,000 to every reward in an environment. The optimal policy is unchanged, and the expected policy gradient is unchanged — but the variance of the estimator you actually compute goes up by four orders of magnitude. Every advan…