Premium Content

Offline Reinforcement Learning

Learning a policy from a fixed dataset

This chapter requires a subscription to access.

What you'll unlock:

  • 1. The Offline RL Problem
  • 2. Why Naive Q-Learning Fails Offline
  • 3. BCQ and BEAR
  • 4. CQL: Conservative Q-Learning
  • 5. IQL: Implicit Q-Learning
  • 6. TD3+BC, AWAC, Cal-QL
  • 7. The D4RL Benchmark
Subscribe to Unlock

Already have an account? Sign in