Multiplier Bootstrap-based Exploration
Runzhe Wan, Haoyu Wei, Branislav Kveton, Rui Song
Abstract
Despite the great interest in the bandit problem, designing efficient algorithms for complex models remains challenging, as there is typically no analytical way to quantify uncertainty. In this paper, we propose Multiplier Bootstrap-based Exploration (MBE), a novel exploration strategy that is applicable to any reward model amenable to weighted loss minimization. We prove both instance-dependent and instance-independent rate-optimal regret bounds for MBE in sub-Gaussian multi-armed bandits. With extensive simulation and real data experiments, we show the generality and adaptivity of MBE.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cb750b9a-3b8b-4b39-afe9-b4791b82c271Cited by top-tier papers3
- Improved Regret of Linear Ensemble SamplingHarin Lee, Min-hwan OhNeurIPS 2024 · 8 citations
- BanditSpec: Adaptive Speculative Decoding via Bandit AlgorithmsYunlong Hou, Fengzhuo Zhang, Cunxiao Du, Xuan Zhang et al.ICML 2025
- Zero-Inflated BanditsHaoyu Wei, Runzhe Wan, Lei Shi, Rui SongICML 2025
Builds on1
Related papers
- Robust Contextual Bandits via BootstrappingQiao Tang, Hong Xie, Yunni Xia, Jia Lee et al.AAAI 2021 · 2 citations
- Forced Exploration in Bandit ProblemsQi Han, Li Zhu, Fei GuoAAAI 2024 · 1 citation
- Revisiting Simple Regret: Fast Rates for Returning a Good ArmYao Zhao, Connor Stephens, Csaba Szepesvári, Kwang-Sung JunICML 2023 · 23 citations
- Optimal Algorithms for Stochastic Multi-Armed Bandits with Heavy Tailed RewardsKyungjae Lee, Hongjun Yang, Sungbin Lim, Songhwai OhNeurIPS 2020 · 34 citations
- Batch Ensemble for Variance Dependent Regret in Stochastic BanditsAsaf B. Cassel, Orin Levy, Yishay MansourAAAI 2025 · 3 citations
