Premium Content

Capstone: GRPO for GSM8K Math Reasoning

Replicating the DeepSeekMath recipe at small scale

This chapter requires a subscription to access.

What you'll unlock:

  • 1. GSM8K and Verifiable Math Rewards
  • 2. Implementing GRPO
  • 3. Training a 1B-Parameter Reasoner
  • 4. Evaluation and Failure Analysis
Subscribe to Unlock

Already have an account? Sign in