Skip to content
← All books
Book · Advanced · 80+ hours

Forging Giants: Training Massive Models from Scratch

The Hidden Math, Intuition, and Engineering Behind 671B-Parameter Models

Master the complete engineering pipeline for training 671B-parameter models. From mathematical foundations through DeepSeek V3 architecture (MLA, MoE), distributed training (DualPipe, FP8), GRPO reasoning, and production deployment.

20Chapters
117Sections
37hReading
5Parts
Curriculum

20 chapters— in publication order.

The capstone

Where the book ends in production.

Chapters 18–20 take everything from Parts I–IV and ship it. Inference, serving, evaluation — the stuff tutorials skip.

Chapter 18·5 sections

Inference Optimisation

Prefill/decode split, KV cache management, speculative decoding, quantisation, and expert load balancing.

Open chapter
Chapter 19·5 sections

Serving Infrastructure

Design and operate a serving system for a 671B MoE model at production scale.

Open chapter
Chapter 20·5 sections

Evaluation, Monitoring, and What's Next

Rigorous evaluation practices, production monitoring, and the future of massive model training.

Open chapter

117 sections. Begin with one.

Chapter 1 — Mathematical Bedrock — is where every reader starts.