Premium Content

RLHF Foundations

Aligning language models with human preferences

This chapter requires a subscription to access.

What you'll unlock:

  • 1. The Alignment Problem
  • 2. Supervised Fine-Tuning Baseline
  • 3. Reward Modeling from Pairwise Preferences
  • 4. PPO for RLHF: The InstructGPT Recipe
  • 5. The KL Constraint and the Reference Model
  • 6. Implementing RLHF on a 125M Model
Subscribe to Unlock

Already have an account? Sign in