Training & Alignment
24 min
Learning From the Log: Off-Policy Policy Learning, From IPS to Counterfactual Risk Minimisation
An unbiased estimate of every policy's value does not give you an unbiased choice of policy. The moment an optimiser searches over importance-weighted estimates, it goes looking for the estimator's noise, and the history of learn…