Skip to content
∑ Praveen T N AI & ML
Concepts Flashcards Writing Editorial AI Feed Graph
Portfolio ↗
Overview Concepts Flashcards Writing Editorial AI Feed Graph Search Back to portfolio ↗
AI & ML/ Writing/Tagged “gae”

Tagged “gae”

1 posts.

Clear
All Model Architecture20 Training & Alignment22 Inference & Serving21 Agents & Orchestration11 Reasoning & Evaluation27 Safety, Security & Governance8 Platforms & Practice22
Training & Alignment 25 min

Why Policy Gradients Need a Baseline: Variance, Trust Regions, and the Road to PPO

Add 1,000 to every reward in an environment. The optimal policy is unchanged, and the expected policy gradient is unchanged — but the variance of the estimator you actually compute goes up by four orders of magnitude. Every advan…

rl policy-gradients reinforce ppo ∑ ◫
The library

1057 concepts, 11,370 flashcards and 131 long-form pieces on AI, NLP, deep learning, LLMs and agentic systems. Free, no sign-up, no paywall.

Sections Concepts Flashcards Writing AI Feed Knowledge Graph
Elsewhere Portfolio Architecture Practice RSS LinkedIn Buy me a coffee Contact Privacy policy Terms of use

Written and maintained by Praveen T N.

© 2026 Praveen T N