Multiplier Bootstrap-based Exploration
Runzhe Wan, Haoyu Wei, Branislav Kveton, Rui Song
2023年份
3被引次数
3顶会引用
摘要
Despite the great interest in the bandit problem, designing efficient algorithms for complex models remains challenging, as there is typically no analytical way to quantify uncertainty. In this paper, we propose Multiplier Bootstrap-based Exploration (MBE), a novel exploration strategy that is applicable to any reward model amenable to weighted loss minimization. We prove both instance-dependent and instance-independent rate-optimal regret bounds for MBE in sub-Gaussian multi-armed bandits. With extensive simulation and real data experiments, we show the generality and adaptivity of MBE.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Improved Regret of Linear Ensemble SamplingHarin Lee, Min-hwan OhNeurIPS 2024 · 被引用 8 次
- BanditSpec: Adaptive Speculative Decoding via Bandit AlgorithmsYunlong Hou, Fengzhuo Zhang, Cunxiao Du, Xuan Zhang 等ICML 2025
- Zero-Inflated BanditsHaoyu Wei, Runzhe Wan, Lei Shi, Rui SongICML 2025
它引用的顶会 Paper1
相关 Paper
- Robust Contextual Bandits via BootstrappingQiao Tang, Hong Xie, Yunni Xia, Jia Lee 等AAAI 2021 · 被引用 2 次
- Forced Exploration in Bandit ProblemsQi Han, Li Zhu, Fei GuoAAAI 2024 · 被引用 1 次
- Revisiting Simple Regret: Fast Rates for Returning a Good ArmYao Zhao, Connor Stephens, Csaba Szepesvári, Kwang-Sung JunICML 2023 · 被引用 23 次
- Optimal Algorithms for Stochastic Multi-Armed Bandits with Heavy Tailed RewardsKyungjae Lee, Hongjun Yang, Sungbin Lim, Songhwai OhNeurIPS 2020 · 被引用 34 次
- Batch Ensemble for Variance Dependent Regret in Stochastic BanditsAsaf B. Cassel, Orin Levy, Yishay MansourAAAI 2025 · 被引用 3 次
