Premium Content

Direct Preference Optimization (DPO)

Closed-form alternatives to RLHF-PPO

This chapter requires a subscription to access.

What you'll unlock:

  • 1. From RLHF to DPO: The Derivation
  • 2. DPO in Practice
  • 3. IPO: Identity Preference Optimization
  • 4. KTO: Kahneman-Tversky Optimization
  • 5. ORPO and SimPO
  • 6. Choosing Among DPO Variants
Subscribe to Unlock

Already have an account? Sign in