Thompson Sampling via Local Uncertainty
Zhendong Wang, Mingyuan Zhou
Abstract
Thompson sampling is an efficient algorithm for sequential decision making, which exploits the posterior uncertainty to address the exploration-exploitation dilemma. There has been significant recent interest in integrating Bayesian neural networks into Thompson sampling. Most of these methods rely on global variable uncertainty for exploration. In this paper, we propose a new probabilistic modeling framework for Thompson sampling, where local latent variable uncertainty is used to sample the mean reward. Variational inference is used to approximate the posterior of the local variable, and semi-implicit structure is further introduced to enhance its expressiveness. Our experimental results on eight contextual bandit benchmark datasets show that Thompson sampling guided by local uncertainty achieves state-of-the-art performance while having low computational complexity.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c0a16dd2-7e30-4f02-b3df-2a23b04f39c5Cited by top-tier papers10
- CARD: Classification and Regression Diffusion ModelsXizewen Han, Huangjie Zheng, Mingyuan ZhouNeurIPS 2022 · 185 citations
- Bayesian Attention Belief NetworksShujian Zhang, Xinjie Fan, Bo Chen, Mingyuan ZhouICML 2021 · 38 citations
- Contextual Dropout: An Efficient Sample-Dependent Dropout ModuleXinjie Fan, Shujian Zhang, Korawat Tanwisuth, Xiaoning Qian et al.ICLR 2021 · 34 citations
- Langevin Monte Carlo for Contextual BanditsPan Xu, Hongkai Zheng, Eric V. Mazumdar, Kamyar Azizzadenesheli et al.ICML 2022 · 34 citations
- Randomized Exploration in Cooperative Multi-Agent Reinforcement LearningHao-Lun Hsu, Weixin Wang, Miroslav Pajic, Pan XuNeurIPS 2024 · 25 citations
Builds on1
Related papers
- Neural Thompson SamplingWeitong Zhang, Dongruo Zhou, Lihong Li, Quanquan GuICLR 2021 · 152 citations
- Deep Bandits Show-Off: Simple and Efficient Exploration with Deep NetworksRong Zhu, Mattia RigottiNeurIPS 2021 · 10 citations
- VITS : Variational Inference Thompson Sampling for contextual banditsPierre Clavier, Tom Huix, Alain Oliviero DurmusICML 2024 · 6 citations
- Thompson Sampling for High-Dimensional Sparse Linear Contextual BanditsSunrit Chakraborty, Saptarshi Roy, Ambuj TewariICML 2023 · 15 citations
- Implicit Generative Modeling for Efficient ExplorationNeale Ratzlaff, Qinxun Bai, Fuxin Li, Wei XuICML 2020 · 15 citations
