Premium ContentThe Policy Gradient Theorem
Direct policy optimization
This chapter requires a subscription to access.
What you'll unlock:
- 1. The Likelihood-Ratio Trick
- 2. Deriving the Policy Gradient Theorem
- 3. REINFORCE
- 4. Variance Reduction with Baselines
Subscribe to UnlockAlready have an account? Sign in