Denoising Diffusion Models from Scratch
From Mathematical Foundations to Image Generation
Master diffusion models from mathematical foundations to image generation. Learn the theory, implement DDPM in PyTorch, and understand modern systems like Stable Diffusion.
18 chapters— in publication order.
Part I · Chapter 00 · 5 sections · 85 min
Essential mathematical background for understanding diffusion models
- 0.1Probability Fundamentals15m
- 0.2Gaussian Distributions Deep Dive20m
- 0.3Information Theory Essentials15m
- 0.4Variational Inference Primer20m
- 0.5Markov Chains and Stochastic Processes15m
Part I · Chapter 01 · 4 sections · 49 min
Context and motivation for diffusion models
- 1.1The Generative Modeling Problem12m
- 1.2Landscape of Generative Models15m
- 1.3The Diffusion Model Intuition12m
- 1.4Historical Context and Key Papers10m
Part II · Chapter 02 · 5 sections · 80 min
Understanding how to systematically destroy data with noise
- 2.1Defining the Forward Process15m
- 2.2The Noise Schedule18m
- 2.3Closed-Form Sampling at Any Timestep20m
- 2.4Properties of the Forward Process12m
- 2.5Implementing the Forward Process15m
Part II · Chapter 03 · 5 sections · 100 min
Learning to denoise - the heart of diffusion models
- 3.1The Reverse Process Goal12m
- 3.2Deriving the True Reverse Distribution25m
- 3.3Parameterizing the Reverse Process18m
- 3.4The Training Objective Derivation25m
- 3.5Score Matching Perspective20m
Part II · Chapter 04 · 4 sections · 60 min
Deep dive into training objectives and their properties
- 4.1The Simplified Loss Explained15m
- 4.2Loss Weighting Strategies18m
- 4.3Connection to Denoising Autoencoders12m
- 4.4Numerical Analysis of the Loss15m
Part III · Chapter 05 · 5 sections · 88 min
The neural network backbone of diffusion models
- 5.1Why U-Net?12m
- 5.2U-Net Building Blocks18m
- 5.3Time Conditioning15m
- 5.4Attention in Diffusion U-Net18m
- 5.5Complete U-Net Implementation25m
Part III · Chapter 06 · 5 sections · 80 min
Assembling all components into a working DDPM
- 6.1The Complete DDPM Class20m
- 6.2Training Loop Implementation18m
- 6.3Sampling Algorithm15m
- 6.4Practical Training Tips15m
- 6.5Debugging and Visualization12m
Part III · Chapter 07 · 5 sections · 70 min
Faster and better generation algorithms
- 7.1Problems with Ancestral Sampling10m
- 7.2DDIM: Deterministic Sampling20m
- 7.3DDIM Implementation15m
- 7.4Advanced Samplers Overview15m
- 7.5Choosing a Sampler10m
Part IV · Chapter 08 · 5 sections · 80 min
Controlling what the model generates
- 8.1Unconditional vs Conditional Models12m
- 8.2Class-Conditional Diffusion15m
- 8.3Classifier Guidance18m
- 8.4Classifier-Free Guidance20m
- 8.5Implementing Classifier-Free Guidance15m
Part IV · Chapter 09 · 4 sections · 63 min
Understanding the path to modern text-to-image models
- 9.1Text Conditioning Overview15m
- 9.2Cross-Attention Mechanism18m
- 9.3CLIP and Contrastive Learning15m
- 9.4From Pixel Space to Latent Space15m
Part V · Chapter 10 · 3 sections · 37 min
Setting up data for training a diffusion model
Start chapter 10Part V · Chapter 11 · 4 sections · 62 min
Complete training setup and execution
- 11.1Model Configuration12m
- 11.2Complete Training Script20m
- 11.3Training Monitoring15m
- 11.4Common Issues and Solutions15m
Part V · Chapter 12 · 4 sections · 57 min
Using the trained model and measuring quality
Start chapter 12Part VI · Chapter 13 · 4 sections · 68 min
Understanding the architecture behind Stable Diffusion
- 13.1The Latent Diffusion Idea15m
- 13.2The VAE Component18m
- 13.3Diffusion in Latent Space15m
- 13.4Stable Diffusion Architecture Overview20m
Part VI · Chapter 14 · 4 sections · 54 min
ControlNet, IP-Adapter, and beyond
- 14.1ControlNet Concept15m
- 14.2Image-to-Image Generation15m
- 14.3IP-Adapter and Image Prompts12m
- 14.4Multi-Modal Conditioning12m
Part VII · Chapter 17 · 4 sections · 44 min
Current research directions
- 17.1Video Generation12m
- 17.23D Generation12m
- 17.3Beyond Images10m
- 17.4Open Problems and Research Directions10m
Where the book lands in practice.
Generation and Evaluation
Using the trained model and measuring quality
Open chapter76 sections. Begin with one.
Chapter 0 — Prerequisites — is where every reader starts.