Premium Content

Distributed Training: DualPipe and the Parallelism Stack

All four parallelism strategies and DeepSeek's DualPipe algorithm that eliminates communication bottlenecks.

This chapter requires a subscription to access.

What you'll unlock:

  • 1. Why One GPU Is Not Enough
  • 2. Data Parallelism (DP)
  • 3. Tensor Parallelism (TP)
  • 4. Pipeline Parallelism and the Bubble Problem
  • 5. DualPipe: DeepSeek's Solution
  • 6. Expert Parallelism and Cross-Node All-to-All
  • 7. Memory Optimisation: No Tensor Parallelism Required
  • 8. Checkpoint Strategy and Fault Tolerance
Subscribe to Unlock

Already have an account? Sign in