Premium Content

The Policy Gradient Theorem

Direct policy optimization

This chapter requires a subscription to access.

What you'll unlock:

  • 1. The Likelihood-Ratio Trick
  • 2. Deriving the Policy Gradient Theorem
  • 3. REINFORCE
  • 4. Variance Reduction with Baselines
Subscribe to Unlock

Already have an account? Sign in