Premium Content

Capstone: RLHF + DPO on a Small Language Model

The full LLM alignment pipeline at toy scale

This chapter requires a subscription to access.

What you'll unlock:

  • 1. SFT on the TL;DR Summarization Dataset
  • 2. Reward Modeling on UltraFeedback
  • 3. PPO RLHF Training
  • 4. DPO Training
  • 5. Comparison and Evaluation
Subscribe to Unlock

Already have an account? Sign in