Premium Content

The 500x Gradient Imbalance

The empirical discovery that motivates the rest of the book: regression gradients exceed classification gradients by 500x on shared parameters.

This chapter requires a subscription to access.

What you'll unlock:

  • 1. Computing Per-Task Gradient Norms
  • 2. Why MSE Gradients Dominate Cross-Entropy
  • 3. Empirical Measurement (n = 4,120 samples)
  • 4. Consequences for Shared Feature Learning
Subscribe to Unlock

Already have an account? Sign in