Premium Content

Reasoning RL: GRPO and Verifiable Rewards

The DeepSeek-R1 generation of RL for reasoning

This chapter requires a subscription to access.

What you'll unlock:

  • 1. RL with Verifiable Rewards (RLVR)
  • 2. GRPO: Group Relative Policy Optimization
  • 3. RLOO and REINFORCE++
  • 4. DAPO: Decoupled Clip and Dynamic Sampling
  • 5. Process Reward Models
  • 6. DeepSeek-R1 Case Study
Subscribe to Unlock

Already have an account? Sign in