RL Foundations

MDPs, value functions, TD learning, policy gradients, actor-critic, TRPO and PPO.

20concepts
140flashcards
146minutes of reading