Premium ContentThe 500x Gradient Imbalance
The empirical discovery that motivates the rest of the book: regression gradients exceed classification gradients by 500x on shared parameters.
This chapter requires a subscription to access.
What you'll unlock:
- 1. Computing Per-Task Gradient Norms
- 2. Why MSE Gradients Dominate Cross-Entropy
- 3. Empirical Measurement (n = 4,120 samples)
- 4. Consequences for Shared Feature Learning
Subscribe to UnlockAlready have an account? Sign in