Premium Content

Proximal Policy Optimization (PPO)

The de-facto workhorse of modern RL

This chapter requires a subscription to access.

What you'll unlock:

  • 1. PPO-Clip
  • 2. PPO-KL (Adaptive KL Penalty)
  • 3. The 37 Implementation Details That Matter
  • 4. PPO for Continuous Control
  • 5. Maskable PPO for Action Constraints
  • 6. Recurrent PPO for POMDPs
Subscribe to Unlock

Already have an account? Sign in