Indice Foundations of Deep RL - lecture series by Pieter Abbeel π Lecture 1 - MDPs, Exact Solution Methods, Max-ent RL Lecture 2 - Deep Q-Learning Lecture 3 - Policy Gradients and Advantage Estimation Lecture 4 - TRPO and PPO Lecture 5 - DDPG and SAC Lecture 6 - Model-based RL Open AI Spinning Up π Intro to Policy Optimization e VPG Trust Region Policy Optimization Proximal Policy Optimization Miscellaneous Note Link utili Generalized Advantage Estimate: Maths and Code π RL β Trust Region Policy Optimization (TRPO) Explained π RL β Trust Region Policy Optimization (TRPO) Part 2 π Implementazioni Baselines π Stable Baselines π Acme π