Premium Content

GRPO: Group Relative Policy Optimisation

Derive GRPO from PPO, eliminate the critic, and implement GRPO from scratch with all reward shaping choices.

This chapter requires a subscription to access.

What you'll unlock:

  • 1. The Critic Bottleneck in PPO
  • 2. GRPO Derivation
  • 3. GRPO Hyperparameters from DeepSeek R1
  • 4. Reward Design for Reasoning
  • 5. GRPO Variants: DAPO, Dr.GRPO, and Olmo 3
  • 6. Implementing GRPO from Scratch
Subscribe to Unlock

Already have an account? Sign in