Tightening Regret Lower and Upper Bounds in Restless Rising Bandits
Cristiano Migali, Marco Mussi, Gianmarco Genalti, Alberto Maria Metelli
Abstract
Restless Multi-Armed Bandits (MABs) are a general framework designed to handle real-world decision-making problems where the expected rewards evolve over time, such as in recommender systems and dynamic pricing. In this work, we investigate from a theoretical standpoint two well-known structured subclasses of restless MABs: the rising and the rising concave settings, where the expected reward of each arm evolves over time following an unknown non-decreasing and a non-decreasing concave function, respectively. By providing a novel methodology of independent interest for general restless bandits, we establish new lower bounds on the expected cumulative regret for both settings. In the rising case, we prove a lower bound of order Ω p T 2 3 q , matching known upper bounds for restless bandits; whereas, in the rising concave case, we derive a lower bound of order Ω p T 3 5 q , proving for the first time that this setting is provably more challenging than stationary MABs. Then, we introduce Rising Concave Budgeted Exploration ( RC-BE p α q ), a new regret minimization algorithm designed for the rising concave MABs. By devising a novel proof technique, we show that the expected cumulative regret of RC-BE p α q is in the order of r O p T 7 11 q . These results collectively make a step towards closing the gap in rising concave MABs, positioning them between stationary and general restless bandit settings in terms of statistical complexity.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3d1963c8-af53-42fe-a45d-273c16a90897Builds on7
- Restless-UCB, an Efficient and Low-complexity Algorithm for Online Restless BanditsSiwei Wang, Longbo Huang, John C. S. LuiNeurIPS 2020 · 58 citations
- Efficient Automatic CASH via Rising BanditsYang Li, Jiawei Jiang, Jinyang Gao, Yingxia Shao et al.AAAI 2020 · 45 citations
- Smooth Non-stationary BanditsSu Jia, Qian Xie, Nathan Kallus, Peter I. FrazierICML 2023 · 14 citations
- Best Model Identification: A Rested Bandit FormulationLeonardo Cella, Massimiliano Pontil, Claudio GentileICML 2021 · 6 citations
- Graph-Triggered Rising BanditsGianmarco Genalti, Marco Mussi, Nicola Gatti, Marcello Restelli et al.ICML 2024 · 6 citations
Related papers
- Non-Stationary Bandits with Auto-Regressive Temporal DependencyQinyi Chen, Negin Golrezaei, Djallel BouneffoufNeurIPS 2023 · 20 citations
- When Demands Evolve Larger and Noisier: Learning and Earning in a Growing EnvironmentFeng Zhu, Zeyu ZhengICML 2020 · 15 citations
- Best Arm Identification for Stochastic Rising BanditsMarco Mussi, Alessandro Montenegro, Francesco Trovò, Marcello Restelli et al.ICML 2024 · 4 citations
- Problem Dependent View on Structured Thresholding Bandit ProblemsJames Cheshire, Pierre Ménard, Alexandra CarpentierICML 2021 · 8 citations
- Stochastic Rising BanditsAlberto Maria Metelli, Francesco Trovò, Matteo Pirola, Marcello RestelliICML 2022 · 1 citation
