Online Sign Identification: Minimization of the Number of Errors in Thresholding Bandits
Reda Ouhamma, Odalric-Ambrym Maillard, Vianney Perchet
Abstract
In the fixed budget thresholding bandit problem, an algorithm sequentially allocates a budgeted number of samples to different distributions. It then predicts whether the mean of each distribution is larger or lower than a given threshold. We introduce a large family of algorithms (containing most existing relevant ones), inspired by the Frank-Wolfe algorithm, and provide a thorough yet generic analysis of their performance. This allowed us to construct new explicit algorithms, for a broad class of problems, whose losses are within a small constant factor of the non-adaptive oracle ones. Quite interestingly, we observed that adaptive methods empirically greatly out-perform non-adaptive oracles, an uncommon behavior in standard online learning settings, such as regret minimization. We explain this surprising phenomenon on an insightful toy problem.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f52ed27b-9b40-4f8c-8cae-6c6f6c859871Cited by top-tier papers1
Ask how each one uses itBuilds on1
Related papers
- Fast Pure Exploration via Frank-WolfePo-An Wang, Ruo-Chun Tzeng, Alexandre ProutièreNeurIPS 2021 · 56 citations
- Adaptive Double-Exploration Tradeoff for Outlier DetectionXiaojin Zhang, Honglei Zhuang, Shengyu Zhang, Yuan ZhouAAAI 2020 · 1 citation
- Adaptive Sampling for Estimating Probability DistributionsShubhanshu Shekhar, Tara Javidi, Mohammad GhavamzadehICML 2020 · 8 citations
- Satisficing Regret Minimization in BanditsQing Feng, Tianyi Ma, Ruihao ZhuICLR 2025 · 1 citation
- Beyond the Best: Distribution Functional Estimation in Infinite-Armed BanditsYifei Wang, Tavor Z. Baharav, Yanjun Han, Jiantao Jiao et al.NeurIPS 2022 · 2 citations
