Premium Content

Scaling Laws and Compute-Optimal Training

Determine the right model size and training token count for a given compute budget using theory and evidence.

This chapter requires a subscription to access.

What you'll unlock:

  • 1. The Chinchilla Scaling Laws
  • 2. MoE Scaling Laws
  • 3. Emergent Abilities
  • 4. Hyperparameter Scaling
  • 5. Predicting Final Loss from Intermediate Checkpoints
  • 6. Inference-Aware Scaling Laws
Subscribe to Unlock

Already have an account? Sign in