Premium ContentProximal Policy Optimization (PPO)
The de-facto workhorse of modern RL
This chapter requires a subscription to access.
What you'll unlock:
- 1. PPO-Clip
- 2. PPO-KL (Adaptive KL Penalty)
- 3. The 37 Implementation Details That Matter
- 4. PPO for Continuous Control
- 5. Maskable PPO for Action Constraints
- 6. Recurrent PPO for POMDPs
Subscribe to UnlockAlready have an account? Sign in