Premium Content

Reward Modeling and RLHF

The full RLHF pipeline from preference data to a trained reward model and PPO training.

This chapter requires a subscription to access.

What you'll unlock:

  • 1. The Alignment Problem
  • 2. The Bradley-Terry Model for Preferences
  • 3. Reward Model Architecture and Training
  • 4. Rule-Based Rewards for Verifiable Tasks
  • 5. Generative Reward Models (LLM-as-Judge)
  • 6. PPO: The Standard RLHF Algorithm
  • 7. DPO: Direct Preference Optimization
Subscribe to Unlock

Already have an account? Sign in