AAAI2021
Meta-Learning Effective Exploration Strategies for Contextual Bandits
Amr Sharaf, Hal Daumé III
被引用 2 次
摘要
In contextual bandits, an algorithm must choose actions given observed contexts, learning from a reward signal that is observed only for the action chosen. This leads to an exploration/exploitation trade-off: the algorithm must balance taking actions it already believes are good with taking new actions to potentially discover better choices. We develop a meta-learning algorithm, M ÊL ÉE, that learns an exploration policy based on simulated, synthetic contextual bandit tasks. M ÊL ÉE uses imitation learning against these simulations to train an exploration policy that can be applied to true contextual bandit tasks at test time. We evaluate M ÊL ÉE on both a natural contextual bandit problem derived from a learning to rank dataset as well as hundreds of simulated contextual bandit problems derived from classification tasks. M ÊL ÉE outperforms seven strong baselines on most of these datasets by leveraging a rich feature representation for learning an exploration strategy.