Premium ContentMulti-Armed Bandits
Exploration and exploitation without state
This chapter requires a subscription to access.
What you'll unlock:
- 1. The k-Armed Bandit Problem
- 2. Action-Value Methods and ε-Greedy
- 3. Upper-Confidence-Bound (UCB)
- 4. Gradient Bandit Algorithms
- 5. Thompson Sampling
- 6. Contextual Bandits and LinUCB
Subscribe to UnlockAlready have an account? Sign in