Forging Giants: Training Massive Models from Scratch
The Hidden Math, Intuition, and Engineering Behind 671B-Parameter Models
Master the complete engineering pipeline for training 671B-parameter models. From mathematical foundations through DeepSeek V3 architecture (MLA, MoE), distributed training (DualPipe, FP8), GRPO reasoning, and production deployment.
20 chapters— in publication order.
Part I · Chapter 01 · 6 sections · 130 min
The specific mathematical tools that recur throughout large-scale training: tensors, SVD, probability, autodiff, optimisation, and numerical precision.
- 1.1Tensors, Matrix Multiply, and Linear Maps20m
- 1.2Eigendecomposition, SVD, and Low-Rank Approximation25m
- 1.3Probability, Entropy, and Cross-Entropy Loss20m
- 1.4Automatic Differentiation and the Chain Rule20m
- 1.5Gradient Descent, Adam, and Optimisation Landscapes25m
- 1.6Numerical Precision Fundamentals20m
Part I · Chapter 02 · 7 sections · 150 min
Build the complete transformer architecture by deriving each component from the problem it solves.
- 2.1The Sequence Modelling Problem15m
- 2.2Scaled Dot-Product Attention: Derivation25m
- 2.3Multi-Head Attention20m
- 2.4Positional Encodings and RoPE25m
- 2.5Feed-Forward Networks as Memory20m
- 2.6Layer Normalisation and Training Stability20m
- 2.7Full Transformer Forward Pass: End-to-End with Shapes25m
Part I · Chapter 03 · 6 sections · 85 min
The full tokenisation pipeline from raw text to token IDs, and the design decisions that affect model quality.
- 3.1Why Subword Tokenisation?10m
- 3.2Byte-Pair Encoding from Scratch20m
- 3.3SentencePiece and Unigram LM Tokenisation15m
- 3.4Vocabulary Size Trade-offs10m
- 3.5Special Tokens and Chat Templates15m
- 3.6Tokenisation Pitfalls15m
Part II · Chapter 04 · 6 sections · 130 min
The KV cache bottleneck at scale and MLA's low-rank compression solution derived completely from scratch.
- 4.1The KV Cache Bottleneck20m
- 4.2Grouped-Query and Multi-Query Attention20m
- 4.3MLA: Low-Rank Joint Compression — Full Derivation30m
- 4.4Decoupled RoPE20m
- 4.5MLA vs GQA vs MHA: Full Comparison15m
- 4.6Implementing MLA in PyTorch25m
Part II · Chapter 05 · 6 sections · 125 min
The MoE architecture from sparse conditional compute, DeepSeekMoE fine-grained expert decomposition, and expert parallelism.
- 5.1Why Mixture-of-Experts?15m
- 5.2The Routing Mechanism25m
- 5.3Fine-Grained Expert Decomposition20m
- 5.4Shared Experts15m
- 5.5Expert Parallelism at Scale25m
- 5.6Implementing DeepSeekMoE25m
Part II · Chapter 06 · 4 sections · 75 min
Why MoE load imbalance is catastrophic, why auxiliary loss hurts quality, and how DeepSeek's bias-term solution resolves the tension.
- 6.1The Routing Collapse Problem15m
- 6.2Auxiliary Loss Approaches and Their Cost20m
- 6.3Bias-Term Load Balancing: The DeepSeek Solution25m
- 6.4Sequence-Level Balance Loss15m
Part II · Chapter 07 · 5 sections · 95 min
Why next-token prediction under-uses the forward pass, DeepSeek's sequential MTP, and speculative decoding at inference.
- 7.1The Single-Token Prediction Bottleneck15m
- 7.2Naive Parallel MTP and Its Failure Mode15m
- 7.3Sequential Causal MTP: DeepSeek's Implementation25m
- 7.4MTP Training Objective and Ablations20m
- 7.5MTP as Speculative Decoding at Inference20m
Part III · Chapter 08 · 7 sections · 120 min
Build a complete data pipeline capable of producing 14.8T high-quality training tokens.
- 8.1The Data Pipeline Architecture20m
- 8.2Deduplication at Scale20m
- 8.3Quality Filtering20m
- 8.4Data Mixing and Domain Weighting15m
- 8.5Synthetic Data: When and How15m
- 8.6Training Data Curriculum15m
- 8.7Data Contamination Detection15m
Part III · Chapter 09 · 6 sections · 120 min
Determine the right model size and training token count for a given compute budget using theory and evidence.
- 9.1The Chinchilla Scaling Laws25m
- 9.2MoE Scaling Laws20m
- 9.3Emergent Abilities15m
- 9.4Hyperparameter Scaling20m
- 9.5Predicting Final Loss from Intermediate Checkpoints15m
- 9.6Inference-Aware Scaling Laws25m
Part III · Chapter 10 · 6 sections · 120 min
Why FP8 training is hard, DeepSeek's fine-grained quantisation and high-precision accumulation solutions.
- 10.1The Case for FP815m
- 10.2Why Naive FP8 Training Fails20m
- 10.3Fine-Grained Quantisation25m
- 10.4High-Precision Accumulation20m
- 10.5What Stays in BF16 and FP3215m
- 10.6Implementing FP8 Training25m
Part III · Chapter 11 · 8 sections · 160 min
All four parallelism strategies and DeepSeek's DualPipe algorithm that eliminates communication bottlenecks.
- 11.1Why One GPU Is Not Enough10m
- 11.2Data Parallelism (DP)20m
- 11.3Tensor Parallelism (TP)20m
- 11.4Pipeline Parallelism and the Bubble Problem25m
- 11.5DualPipe: DeepSeek's Solution30m
- 11.6Expert Parallelism and Cross-Node All-to-All20m
- 11.7Memory Optimisation: No Tensor Parallelism Required20m
- 11.8Checkpoint Strategy and Fault Tolerance15m
Part III · Chapter 12 · 4 sections · 65 min
Why extending context is non-trivial, and the YaRN technique used by DeepSeek V3 to reach 128K tokens.
- 12.1Why Context Extension Is Hard15m
- 12.2NTK-Aware Scaling15m
- 12.3YaRN: Frequency-Domain Interpolation25m
- 12.4Evaluation at Long Context10m
Part IV · Chapter 13 · 5 sections · 80 min
SFT as a formatting and style adapter on top of a knowledge-rich base model.
- 13.1What SFT Actually Does15m
- 13.2SFT Data Collection15m
- 13.3Chat Template and Formatting15m
- 13.4SFT Training Configuration20m
- 13.5Catastrophic Forgetting and Mitigation15m
Part IV · Chapter 14 · 7 sections · 140 min
The full RLHF pipeline from preference data to a trained reward model and PPO training.
- 14.1The Alignment Problem10m
- 14.2The Bradley-Terry Model for Preferences20m
- 14.3Reward Model Architecture and Training20m
- 14.4Rule-Based Rewards for Verifiable Tasks15m
- 14.5Generative Reward Models (LLM-as-Judge)15m
- 14.6PPO: The Standard RLHF Algorithm30m
- 14.7DPO: Direct Preference Optimization30m
Part IV · Chapter 15 · 6 sections · 130 min
Derive GRPO from PPO, eliminate the critic, and implement GRPO from scratch with all reward shaping choices.
- 15.1The Critic Bottleneck in PPO15m
- 15.2GRPO Derivation30m
- 15.3GRPO Hyperparameters from DeepSeek R115m
- 15.4Reward Design for Reasoning20m
- 15.5GRPO Variants: DAPO, Dr.GRPO, and Olmo 320m
- 15.6Implementing GRPO from Scratch30m
Part IV · Chapter 16 · 6 sections · 80 min
The R1-Zero experiment — its hypothesis, results, emergent phenomena, and what it reveals about LLM reasoning.
- 16.1The R1-Zero Hypothesis15m
- 16.2Experimental Setup10m
- 16.3The Aha Moment15m
- 16.4Quantitative Results15m
- 16.5Failure Modes of R1-Zero10m
- 16.6What R1-Zero Teaches Us15m
Part IV · Chapter 17 · 7 sections · 105 min
The multi-stage pipeline from base model to production reasoning model, including distillation to smaller models.
- 17.1Pipeline Overview15m
- 17.2Cold-Start SFT: Why a Few Thousand Examples Help15m
- 17.3Stage 2: Reasoning-Focused GRPO15m
- 17.4Stage 3: Rejection Sampling for SFT Data15m
- 17.5Stage 4: Alignment GRPO15m
- 17.6Distillation to Smaller Models20m
- 17.7Distilling Reasoning into DeepSeek V3 Chat10m
Part V · Chapter 18 · 5 sections · 100 min
Prefill/decode split, KV cache management, speculative decoding, quantisation, and expert load balancing.
- 18.1Two Very Different Problems: Prefill vs Decode20m
- 18.2KV Cache Management and PagedAttention25m
- 18.3Speculative Decoding with MTP20m
- 18.4Post-Training Quantisation20m
- 18.5Expert Load Balancing at Inference15m
Part V · Chapter 19 · 5 sections · 95 min
Design and operate a serving system for a 671B MoE model at production scale.
- 19.1The DeepSeek Production Architecture20m
- 19.2Inference Framework Comparison15m
- 19.3Continuous Batching and Request Scheduling20m
- 19.4Multi-Node Inference and Network Requirements20m
- 19.5Cost Modelling and Optimisation20m
Part V · Chapter 20 · 5 sections · 90 min
Rigorous evaluation practices, production monitoring, and the future of massive model training.
- 20.1Benchmark Taxonomy15m
- 20.2Evaluation Contamination Detection15m
- 20.3Production Monitoring20m
- 20.4The Future: What DeepSeek's Work Reveals15m
- 20.5Reproducing DeepSeek on a Budget: A Practical Roadmap25m
Where the book ends in production.
Chapters 18–20 take everything from Parts I–IV and ship it. Inference, serving, evaluation — the stuff tutorials skip.
Inference Optimisation
Prefill/decode split, KV cache management, speculative decoding, quantisation, and expert load balancing.
Open chapterServing Infrastructure
Design and operate a serving system for a 671B MoE model at production scale.
Open chapterEvaluation, Monitoring, and What's Next
Rigorous evaluation practices, production monitoring, and the future of massive model training.
Open chapter117 sections. Begin with one.
Chapter 1 — Mathematical Bedrock — is where every reader starts.