Premium ContentCapstone: RLHF + DPO on a Small Language Model
The full LLM alignment pipeline at toy scale
This chapter requires a subscription to access.
What you'll unlock:
- 1. SFT on the TL;DR Summarization Dataset
- 2. Reward Modeling on UltraFeedback
- 3. PPO RLHF Training
- 4. DPO Training
- 5. Comparison and Evaluation
Subscribe to UnlockAlready have an account? Sign in