Skip to content
∑ Praveen T N AI & ML
Concepts Flashcards Writing Editorial AI Feed Graph
Portfolio ↗
Overview Concepts Flashcards Writing Editorial AI Feed Graph Search Back to portfolio ↗
AI & ML/ Writing/Tagged “policy-gradients”

Tagged “policy-gradients”

2 posts.

Clear
All Model Architecture20 Training & Alignment22 Inference & Serving21 Agents & Orchestration11 Reasoning & Evaluation27 Safety, Security & Governance8 Platforms & Practice22
Training & Alignment 24 min

Learning From the Log: Off-Policy Policy Learning, From IPS to Counterfactual Risk Minimisation

An unbiased estimate of every policy's value does not give you an unbiased choice of policy. The moment an optimiser searches over importance-weighted estimates, it goes looking for the estimator's noise, and the history of learn…

policy-learning-and-ope off-policy-evaluation counterfactual importance-sampling ∑ ◫
Training & Alignment 25 min

Why Policy Gradients Need a Baseline: Variance, Trust Regions, and the Road to PPO

Add 1,000 to every reward in an environment. The optimal policy is unchanged, and the expected policy gradient is unchanged — but the variance of the estimator you actually compute goes up by four orders of magnitude. Every advan…

rl policy-gradients reinforce ppo ∑ ◫
The library

1057 concepts, 11,370 flashcards and 131 long-form pieces on AI, NLP, deep learning, LLMs and agentic systems. Free, no sign-up, no paywall.

Sections Concepts Flashcards Writing AI Feed Knowledge Graph
Elsewhere Portfolio Architecture Practice RSS LinkedIn Buy me a coffee Contact Privacy policy Terms of use

Written and maintained by Praveen T N.

© 2026 Praveen T N