Thompson Sampling via Local Uncertainty
Zhendong Wang, Mingyuan Zhou
摘要
Thompson sampling is an efficient algorithm for sequential decision making, which exploits the posterior uncertainty to address the exploration-exploitation dilemma. There has been significant recent interest in integrating Bayesian neural networks into Thompson sampling. Most of these methods rely on global variable uncertainty for exploration. In this paper, we propose a new probabilistic modeling framework for Thompson sampling, where local latent variable uncertainty is used to sample the mean reward. Variational inference is used to approximate the posterior of the local variable, and semi-implicit structure is further introduced to enhance its expressiveness. Our experimental results on eight contextual bandit benchmark datasets show that Thompson sampling guided by local uncertainty achieves state-of-the-art performance while having low computational complexity.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- CARD: Classification and Regression Diffusion ModelsXizewen Han, Huangjie Zheng, Mingyuan ZhouNeurIPS 2022 · 被引用 185 次
- Bayesian Attention Belief NetworksShujian Zhang, Xinjie Fan, Bo Chen, Mingyuan ZhouICML 2021 · 被引用 38 次
- Contextual Dropout: An Efficient Sample-Dependent Dropout ModuleXinjie Fan, Shujian Zhang, Korawat Tanwisuth, Xiaoning Qian 等ICLR 2021 · 被引用 34 次
- Langevin Monte Carlo for Contextual BanditsPan Xu, Hongkai Zheng, Eric V. Mazumdar, Kamyar Azizzadenesheli 等ICML 2022 · 被引用 34 次
- Randomized Exploration in Cooperative Multi-Agent Reinforcement LearningHao-Lun Hsu, Weixin Wang, Miroslav Pajic, Pan XuNeurIPS 2024 · 被引用 25 次
它引用的顶会 Paper1
相关 Paper
- Neural Thompson SamplingWeitong Zhang, Dongruo Zhou, Lihong Li, Quanquan GuICLR 2021 · 被引用 152 次
- Deep Bandits Show-Off: Simple and Efficient Exploration with Deep NetworksRong Zhu, Mattia RigottiNeurIPS 2021 · 被引用 10 次
- VITS : Variational Inference Thompson Sampling for contextual banditsPierre Clavier, Tom Huix, Alain Oliviero DurmusICML 2024 · 被引用 6 次
- Thompson Sampling for High-Dimensional Sparse Linear Contextual BanditsSunrit Chakraborty, Saptarshi Roy, Ambuj TewariICML 2023 · 被引用 15 次
- Implicit Generative Modeling for Efficient ExplorationNeale Ratzlaff, Qinxun Bai, Fuxin Li, Wei XuICML 2020 · 被引用 15 次
