Training & Alignment
25 min
The Four Faces of KL Divergence: Mode-Seeking, Mode-Covering, and Why Your Estimator Went Negative
You add a KL penalty to an RLHF objective, log it, and it prints minus 0.03. KL divergence is provably non-negative, and nothing is broken. One formula does four different jobs in modern machine learning, and almost every confusi…