Advanced Reinforcement Learning
Volume II — Alignment, Multi-Agent, and Frontier
The frontier half of the RL curriculum: imitation and offline RL (CQL, IQL, Decision Transformer), hierarchical and meta-RL, multi-agent (MADDPG, QMIX, MAPPO, PSRO), RL for language models (RLHF, DPO, GRPO, DAPO, DeepSeek-R1), distributed engineering, and six capstone projects. Volume 2 of 2 — assumes the foundations covered in Volume 1.
19 chapters— in publication order.
Part I · Chapter 01 · 5 sections · 106 min
Learning from demonstrations and inferring rewards
- 1.1Behavior Cloning18m
- 1.2Distribution Shift and the BC Failure Mode18m
- 1.3DAgger: Dataset Aggregation20m
- 1.4Maximum-Entropy Inverse RL25m
- 1.5GAIL and AIRL25m
Part I · Chapter 02 · 7 sections · 157 min
Learning a policy from a fixed dataset
- 2.1The Offline RL Problem20m
- 2.2Why Naive Q-Learning Fails Offline22m
- 2.3BCQ and BEAR22m
- 2.4CQL: Conservative Q-Learning28m
- 2.5IQL: Implicit Q-Learning25m
- 2.6TD3+BC, AWAC, Cal-QL22m
- 2.7The D4RL Benchmark18m
Part I · Chapter 03 · 5 sections · 107 min
RL as supervised sequence prediction
- 3.1RL as Sequence Modeling18m
- 3.2Decision Transformer25m
- 3.3Trajectory Transformer22m
- 3.4Online Decision Transformer20m
- 3.5Multi-Game DT and Gato22m
Part II · Chapter 04 · 5 sections · 106 min
Structure across temporal scales
- 4.1The Options Framework22m
- 4.2FeUdal Networks22m
- 4.3HIRO: Hierarchical Off-Policy RL20m
- 4.4Goal-Conditioned RL20m
- 4.5Hindsight Experience Replay (HER)22m
Part II · Chapter 05 · 4 sections · 87 min
Learning to learn
- 5.1The Meta-RL Problem18m
- 5.2MAML for Reinforcement Learning25m
- 5.3RL² and In-Context RL22m
- 5.4PEARL: Probabilistic Embeddings for Meta-RL22m
Part II · Chapter 06 · 6 sections · 134 min
Cooperation, competition, and self-play
- 6.1Game-Theoretic Foundations22m
- 6.2Independent Learners and Non-Stationarity18m
- 6.3VDN and QMIX: Value Decomposition25m
- 6.4MADDPG and COMA22m
- 6.5MAPPO: PPO for Cooperative MARL22m
- 6.6Self-Play and PSRO: AlphaStar and OpenAI Five25m
Part III · Chapter 07 · 6 sections · 139 min
Aligning language models with human preferences
- 7.1The Alignment Problem18m
- 7.2Supervised Fine-Tuning Baseline18m
- 7.3Reward Modeling from Pairwise Preferences25m
- 7.4PPO for RLHF: The InstructGPT Recipe28m
- 7.5The KL Constraint and the Reference Model20m
- 7.6Implementing RLHF on a 125M Model30m
Part III · Chapter 08 · 6 sections · 121 min
Closed-form alternatives to RLHF-PPO
- 8.1From RLHF to DPO: The Derivation25m
- 8.2DPO in Practice22m
- 8.3IPO: Identity Preference Optimization18m
- 8.4KTO: Kahneman-Tversky Optimization18m
- 8.5ORPO and SimPO20m
- 8.6Choosing Among DPO Variants18m
Part III · Chapter 09 · 6 sections · 142 min
The DeepSeek-R1 generation of RL for reasoning
- 9.1RL with Verifiable Rewards (RLVR)22m
- 9.2GRPO: Group Relative Policy Optimization28m
- 9.3RLOO and REINFORCE++22m
- 9.4DAPO: Decoupled Clip and Dynamic Sampling20m
- 9.5Process Reward Models22m
- 9.6DeepSeek-R1 Case Study28m
Part III · Chapter 10 · 4 sections · 91 min
Synthetic preferences and self-improvement
- 10.1RLAIF: Synthetic Preference Data22m
- 10.2Constitutional AI25m
- 10.3Self-Rewarding Language Models20m
- 10.4Scalable Oversight: Debate and Weak-to-Strong24m
Part IV · Chapter 11 · 5 sections · 99 min
From notebook to thousands of GPUs
- 11.1Vectorized Environments18m
- 11.2Replay Buffer Engineering20m
- 11.3Mixed Precision and Memory Optimization18m
- 11.4Distributed RL: Ape-X, R2D2, SEED, Sebulba25m
- 11.5Reproducibility and the Deep-RL Crisis18m
Part IV · Chapter 12 · 5 sections · 104 min
Beyond simulators
- 12.1Sim2Real and Domain Randomization22m
- 12.2Robotics with Isaac Lab22m
- 12.3Recommender Systems and Ads20m
- 12.4Control and Operations Research18m
- 12.5Scientific RL: AlphaFold and Beyond22m
Part IV · Chapter 13 · 4 sections · 76 min
Measuring progress honestly
- 13.1The Atari ALE: Strengths and Pitfalls18m
- 13.2MuJoCo and Brax18m
- 13.3Procgen and Generalization20m
- 13.4Crafter and MineRL20m
Part V · Chapter 14 · 4 sections · 95 min
End-to-end discrete-action deep RL
- 14.1Setup and Atari Preprocessing18m
- 14.2Building the Rainbow Agent30m
- 14.3Training and Hyperparameter Tuning25m
- 14.4Evaluation and Ablation Study22m
Part V · Chapter 15 · 4 sections · 86 min
Continuous control mastery
- 15.1MuJoCo Humanoid Environment18m
- 15.2Building the SAC Agent28m
- 15.3Training and Curriculum22m
- 15.4Analysis and Locomotion Visualization18m
Part V · Chapter 16 · 4 sections · 90 min
MCTS plus self-play in PyTorch
- 16.1Connect Four: Rules and State Representation15m
- 16.2Building the Policy-Value Network22m
- 16.3MCTS with Neural Priors28m
- 16.4Self-Play Training Loop25m
Part V · Chapter 17 · 4 sections · 95 min
Learned world models, end-to-end
- 17.1Crafter: Long-Horizon Survival18m
- 17.2The DreamerV3 World Model30m
- 17.3Actor-Critic in Imagination25m
- 17.4Training, Evaluation, and Ablations22m
Part V · Chapter 18 · 5 sections · 114 min
The full LLM alignment pipeline at toy scale
- 18.1SFT on the TL;DR Summarization Dataset20m
- 18.2Reward Modeling on UltraFeedback22m
- 18.3PPO RLHF Training28m
- 18.4DPO Training22m
- 18.5Comparison and Evaluation22m
Part V · Chapter 19 · 4 sections · 93 min
Replicating the DeepSeekMath recipe at small scale
- 19.1GSM8K and Verifiable Math Rewards18m
- 19.2Implementing GRPO28m
- 19.3Training a 1B-Parameter Reasoner25m
- 19.4Evaluation and Failure Analysis22m
Where the book lands in practice.
Capstone: RLHF + DPO on a Small Language Model
The full LLM alignment pipeline at toy scale
Open chapterCapstone: GRPO for GSM8K Math Reasoning
Replicating the DeepSeekMath recipe at small scale
Open chapter93 sections. Begin with one.
Chapter 1 — Imitation and Inverse RL — is where every reader starts.