Premium Content

Trust-Region Methods

Stable policy updates

This chapter requires a subscription to access.

What you'll unlock:

  • 1. Why Naive Policy Gradient Fails
  • 2. The Natural Policy Gradient
  • 3. The Fisher Information Matrix
  • 4. TRPO Derivation
  • 5. Implementing TRPO
Subscribe to Unlock

Already have an account? Sign in