Meta-Learning Effective Exploration Strategies for Contextual Bandits
Amr Sharaf, Hal Daumé III
Abstract
In contextual bandits, an algorithm must choose actions given observed contexts, learning from a reward signal that is observed only for the action chosen. This leads to an exploration/exploitation trade-off: the algorithm must balance taking actions it already believes are good with taking new actions to potentially discover better choices. We develop a meta-learning algorithm, M ÊL ÉE, that learns an exploration policy based on simulated, synthetic contextual bandit tasks. M ÊL ÉE uses imitation learning against these simulations to train an exploration policy that can be applied to true contextual bandit tasks at test time. We evaluate M ÊL ÉE on both a natural contextual bandit problem derived from a learning to rank dataset as well as hundreds of simulated contextual bandit problems derived from classification tasks. M ÊL ÉE outperforms seven strong baselines on most of these datasets by leveraging a rich feature representation for learning an exploration strategy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c2b178e4-a024-4af0-a48f-09a33a963ff3Cited by top-tier papers1
Ask how each one uses itRelated papers
- A Classification View on Meta Learning BanditsMirco Mutti, Jeongyeol Kwon, Shie Mannor, Aviv TamarICML 2025
- Meta-Learning with Neural Bandit SchedulerYunzhe Qi, Yikun Ban, Tianxin Wei, Jiaru Zou et al.NeurIPS 2023 · 14 citations
- EE-Net: Exploitation-Exploration Neural Networks in Contextual BanditsYikun Ban, Yuchen Yan, Arindam Banerjee, Jingrui HeICLR 2022 · 62 citations
- Information-theoretic Task Selection for Meta-Reinforcement LearningRicardo Luna Gutiérrez, Matteo LeonettiNeurIPS 2020 · 24 citations
- Reward-Biased Maximum Likelihood Estimation for Neural Contextual Bandits: A Distributional Learning PerspectiveYu-Heng Hung, Ping-Chun HsiehAAAI 2023 · 2 citations
