Notes on Deep RL, including the different classes of algorithms, implementation details in code, and applications to robotics, games, etc. Heavily based on the Deep RL course provided by Hugging Face.
Table of Contents
- The Ingredients of RL
- Markov Decision Processes
- Dynamic Programming for MDPs
- POMDPs
- Q-Learning
- Policy Gradient
- Actor-Critic Methods
- Proximal Policy Optimization (PPO)
- Advanced Policy Optimization
- Bandits and Exploration
- Offline RL
- Multi-Agent Reinforcement Learning
Applications
Reinforcement Learning from Human Feedback (RLHF)