Reinforcement Learning from Scratch with PyTorch
Volume I — Foundations and Deep RL
Master reinforcement learning from multi-armed bandits to DreamerV3 and MuZero. Derive every algorithm, then implement it in PyTorch: tabular methods, DQN/Rainbow, PPO, SAC, TD3, model-based RL, MCTS, AlphaZero. The classical-through-frontier foundation (Volume 1 of 2).
23 chapters— in publication order.
Part I · Chapter 00 · 4 sections · 52 min
Tools and frameworks for hands-on RL
- 0.1Python Environment for RL10m
- 0.2PyTorch, Gymnasium, and MuJoCo15m
- 0.3CleanRL, Stable Baselines3, and TorchRL12m
- 0.4GPU and Vectorized Environments15m
Part I · Chapter 01 · 5 sections · 83 min
Framing learning from interaction
- 1.1The Agent-Environment Loop18m
- 1.2The Reward Hypothesis15m
- 1.3RL vs Supervised vs Unsupervised15m
- 1.4A Short History: TD-Gammon to ChatGPT20m
- 1.5Two Cultures: Tabular Theory and Deep RL Practice15m
Part I · Chapter 02 · 6 sections · 120 min
Exploration and exploitation without state
- 2.1The k-Armed Bandit Problem18m
- 2.2Action-Value Methods and ε-Greedy20m
- 2.3Upper-Confidence-Bound (UCB)22m
- 2.4Gradient Bandit Algorithms18m
- 2.5Thompson Sampling20m
- 2.6Contextual Bandits and LinUCB22m
Part I · Chapter 03 · 6 sections · 127 min
The formal language of RL
- 3.1States, Actions, Transitions, Rewards18m
- 3.2Returns, Discounting, and Episodes18m
- 3.3Policies and Value Functions22m
- 3.4Bellman Expectation Equations22m
- 3.5Bellman Optimality Equations22m
- 3.6POMDPs and the Contraction-Mapping Proof25m
Part II · Chapter 04 · 5 sections · 92 min
Planning with a known model
- 4.1Iterative Policy Evaluation20m
- 4.2Policy Improvement and Policy Iteration22m
- 4.3Value Iteration20m
- 4.4Generalized Policy Iteration15m
- 4.5Asynchronous DP and Sweep Strategies15m
Part II · Chapter 05 · 5 sections · 100 min
Learning from sample episodes
- 5.1First-Visit and Every-Visit MC20m
- 5.2On-Policy MC Control20m
- 5.3Off-Policy MC with Importance Sampling25m
- 5.4Weighted Importance Sampling20m
- 5.5Preview: Monte Carlo Tree Search15m
Part II · Chapter 06 · 7 sections · 149 min
Bootstrapping from one step ahead
- 6.1TD(0) Prediction20m
- 6.2SARSA: On-Policy TD Control22m
- 6.3Q-Learning: Off-Policy TD Control25m
- 6.4Expected SARSA15m
- 6.5Double Q-Learning and Maximization Bias20m
- 6.6n-Step TD Methods22m
- 6.7TD(λ) and Eligibility Traces25m
Part II · Chapter 07 · 4 sections · 75 min
Unifying model-free and model-based
- 7.1Tabular Dyna-Q22m
- 7.2Dyna-Q+ for Non-Stationary Worlds18m
- 7.3Prioritized Sweeping20m
- 7.4Trajectory Sampling vs Uniform Updates15m
Part III · Chapter 08 · 5 sections · 101 min
Generalizing across states
- 8.1Why Tables Do Not Scale15m
- 8.2Linear Value Functions22m
- 8.3Tile Coding and Fourier Bases20m
- 8.4Semi-Gradient TD22m
- 8.5The Deadly Triad22m
Part III · Chapter 09 · 4 sections · 89 min
Direct policy optimization
- 9.1The Likelihood-Ratio Trick20m
- 9.2Deriving the Policy Gradient Theorem25m
- 9.3REINFORCE22m
- 9.4Variance Reduction with Baselines22m
Part IV · Chapter 10 · 6 sections · 136 min
Neural networks meet Q-learning
- 10.1The DQN Architecture22m
- 10.2Replay Buffers18m
- 10.3Target Networks18m
- 10.4Implementing DQN for CartPole28m
- 10.5DQN for Atari: CNN Encoder and Preprocessing30m
- 10.6Common DQN Pitfalls and Debugging20m
Part IV · Chapter 11 · 6 sections · 116 min
Six tricks that compound into Rainbow
- 11.1Double DQN18m
- 11.2Dueling DQN18m
- 11.3Prioritized Experience Replay22m
- 11.4NoisyNets18m
- 11.5Multi-Step Targets15m
- 11.6Rainbow: Integrating Six Improvements25m
Part IV · Chapter 12 · 5 sections · 105 min
Learning the distribution over returns
- 12.1Why Distributions Over Returns?18m
- 12.2C51: Categorical Distributional RL25m
- 12.3Quantile Regression DQN (QR-DQN)22m
- 12.4Implicit Quantile Networks (IQN)22m
- 12.5Fully Parameterized Quantile Functions (FQF)18m
Part V · Chapter 13 · 5 sections · 109 min
Combining policy and value
- 13.1From REINFORCE to Actor-Critic20m
- 13.2Advantage Estimation20m
- 13.3A2C: Synchronous Advantage Actor-Critic22m
- 13.4A3C: Asynchronous Actor-Critic22m
- 13.5Generalized Advantage Estimation (GAE)25m
Part V · Chapter 14 · 5 sections · 110 min
Stable policy updates
- 14.1Why Naive Policy Gradient Fails15m
- 14.2The Natural Policy Gradient22m
- 14.3The Fisher Information Matrix20m
- 14.4TRPO Derivation28m
- 14.5Implementing TRPO25m
Part V · Chapter 15 · 6 sections · 132 min
The de-facto workhorse of modern RL
- 15.1PPO-Clip22m
- 15.2PPO-KL (Adaptive KL Penalty)18m
- 15.3The 37 Implementation Details That Matter30m
- 15.4PPO for Continuous Control22m
- 15.5Maskable PPO for Action Constraints18m
- 15.6Recurrent PPO for POMDPs22m
Part V · Chapter 16 · 4 sections · 85 min
Scaling PPO and beyond
- 16.1IMPALA and V-Trace25m
- 16.2Phasic Policy Gradient (PPG)18m
- 16.3APPO and the Sebulba Architecture20m
- 16.4Massively Parallel PPO (Isaac Gym Style)22m
Part VI · Chapter 17 · 5 sections · 112 min
Deterministic policy gradients for continuous control
- 17.1The Deterministic Policy Gradient22m
- 17.2DDPG25m
- 17.3TD3: Twin Critics and Delayed Updates25m
- 17.4Target Policy Smoothing15m
- 17.5TD3 on MuJoCo: Hopper, HalfCheetah, Ant25m
Part VI · Chapter 18 · 6 sections · 135 min
Maximum-entropy RL
- 18.1The Maximum-Entropy Framework22m
- 18.2SAC Derivation28m
- 18.3Automatic Temperature Tuning20m
- 18.4Discrete-Action SAC18m
- 18.5Modern Variants: TQC, CrossQ, DroQ, REDQ22m
- 18.6SAC on Humanoid25m
Part VII · Chapter 19 · 6 sections · 129 min
Beyond ε-greedy
- 19.1Why Classical Exploration Fails at Scale18m
- 19.2Count-Based and Pseudo-Count Methods22m
- 19.3Intrinsic Curiosity Module (ICM)22m
- 19.4Random Network Distillation (RND)22m
- 19.5Never Give Up (NGU) and Agent5725m
- 19.6BYOL-Explore and Random Latent Exploration20m
Part VIII · Chapter 20 · 6 sections · 145 min
Learning to predict and plan
- 20.1Why Model-Based RL?18m
- 20.2PETS: Probabilistic Ensembles with Trajectory Sampling25m
- 20.3MBPO: Model-Based Policy Optimization22m
- 20.4World Models (Ha and Schmidhuber)22m
- 20.5Dreamer V1 and V228m
- 20.6DreamerV3: Mastering Diverse Domains30m
Part VIII · Chapter 21 · 5 sections · 117 min
From MCTS to AlphaZero
- 21.1Monte Carlo Tree Search25m
- 21.2UCT: Upper Confidence Trees20m
- 21.3AlphaGo: Policy and Value Networks25m
- 21.4AlphaGo Zero: Pure Self-Play22m
- 21.5AlphaZero: One Algorithm for Three Games25m
Part VIII · Chapter 22 · 5 sections · 116 min
Planning without a given model
- 22.1Learning the Model From Scratch20m
- 22.2The MuZero Algorithm30m
- 22.3EfficientZero: Sample-Efficient MuZero22m
- 22.4Sampled and Gumbel MuZero22m
- 22.5UniZero and Modern Planning Agents22m
121 sections. Begin with one.
Chapter 0 — Development Environment — is where every reader starts.