Premium ContentTrust-Region Methods
Stable policy updates
This chapter requires a subscription to access.
What you'll unlock:
- 1. Why Naive Policy Gradient Fails
- 2. The Natural Policy Gradient
- 3. The Fisher Information Matrix
- 4. TRPO Derivation
- 5. Implementing TRPO
Subscribe to UnlockAlready have an account? Sign in