Notes on Deep RL, including the different classes of algorithms, implementation details in code, and applications to robotics, games, etc. Heavily based on the Deep RL course provided by Hugging Face.

Table of Contents

  1. The Ingredients of RL
  2. Markov Decision Processes
    1. Dynamic Programming for MDPs
    2. POMDPs
  3. Q-Learning
  4. Policy Gradient
    1. Actor-Critic Methods
    2. Proximal Policy Optimization (PPO)
    3. Advanced Policy Optimization
  5. Bandits and Exploration
  6. Offline RL
  7. Multi-Agent Reinforcement Learning

Applications

Reinforcement Learning from Human Feedback (RLHF)

Algorithms (from scratch)

Implementation notes with the update rule and short code for each algorithm: Algorithms index.

  • Bandits: Epsilon-Greedy, UCB, Gradient Bandit
  • Dynamic programming: Policy Evaluation, Value Iteration, Policy Iteration
  • Monte Carlo: First-Visit Monte Carlo, Importance Sampling
  • Temporal difference: TD(0), SARSA, Q-Learning Algorithm, Double Q-Learning, N-Step TD, Eligibility Traces
  • Deep RL: DQN, Prioritized Experience Replay, Dueling Networks, Rainbow DQN, Distributional RL
  • Policy gradient: REINFORCE, Advantage Actor-Critic, GAE, PPO Clipped Objective, TRPO, Deterministic Policy Gradient
  • Model-based: Dyna-Q, Monte Carlo Tree Search, Prioritized Sweeping
  • Post-training: RLHF and preference optimization (DPO, GRPO, Reward Model Training)

References

Hugging Face Deep RL Course

CS 285 Berkeley

3 items under this folder.