Skip to content
∑ Praveen T N Learning Library
Overview Concepts Flashcards Writing AI Feed Graph
Portfolio ↗
Overview Concepts Flashcards Writing AI Feed Graph Search Back to portfolio ↗
Library/ Writing/Tagged “policy-gradients”

Tagged “policy-gradients”

1 posts.

Clear
All Model Architecture14 Training & Alignment14 Inference & Serving12 Agents & Orchestration10 Reasoning & Evaluation8 Safety, Security & Governance3 Platforms & Practice12
Training & Alignment 25 min

Why Policy Gradients Need a Baseline: Variance, Trust Regions, and the Road to PPO

Add 1,000 to every reward in an environment. The optimal policy is unchanged, and the expected policy gradient is unchanged — but the variance of the estimator you actually compute goes up by four orders of magnitude. Every advan…

rl policy-gradients reinforce ppo ∑ ◫
The library

500 concepts, 3,265 flashcards and 73 long-form pieces on AI, NLP, deep learning, LLMs and agentic systems. Free, no sign-up, no paywall.

Sections Concepts Flashcards Writing AI Feed Knowledge Graph
Elsewhere Portfolio RSS LinkedIn Buy me a coffee

Written and maintained by Praveen T N.

© 2026 Praveen T N