p-Mean Regret for Stochastic Bandits
Anand Krishna, Philips George John, Adarsh Barik, Vincent Y. F. Tan
Abstract
In this work, we extend the concept of the p-mean welfare objective from social choice theory (Moulin 2004) to study pmean regret in stochastic multi-armed bandit problems. The p-mean regret, defined as the difference between the optimal mean among the arms and the p-mean of the expected rewards, offers a flexible framework for evaluating bandit algorithms, enabling algorithm designers to balance fairness and efficiency by adjusting the parameter p. Our framework encompasses both average cumulative regret and Nash regret as special cases. We introduce a simple, unified UCBbased algorithm (EXPLORE-THEN-UCB) that achieves novel p-mean regret bounds. Our algorithm consists of two phases: a carefully calibrated uniform exploration phase to initialize sample means, followed by the UCB1 algorithm of Auer, Cesa-Bianchi, and Fischer (2002) . Under mild assumptions, we prove that our algorithm achieves a p-mean regret bound of Õ k T 1 2|p| for all p ≤ -1, where k represents the number of arms and T the time horizon. When -1 < p < 0, we achieve a regret bound of Õ k 1.5 T 1 2 . For the range 0 < p ≤ 1, we achieve a p-mean regret scaling as Õ k T , which matches the previously established lower bound up to logarithmic factors (Auer et al. 1995) . This result stems from the fact that the p-mean regret of any algorithm is at least its average cumulative regret for p ≤ 1. In the case of Nash regret (the limit as p approaches zero), our unified approach differs from prior work (Barman et al. 2023) , which requires a new Nash Confidence Bound algorithm. Notably, we achieve the same regret bound up to constant factors using our more general method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cbb9178c-5fe0-4078-a57f-d988d3ffa59dCited by top-tier papers3
- DP-NCB: Privacy Preserving Fair BanditsDhruv Sarkar, Nishant Pandey, Sayak Ray ChowdhuryAAAI 2026 · 2 citations
- Improved Algorithms for Nash Welfare in Linear BanditsDhruv Sarkar, Nishant Pandey, Sayak Ray ChowdhuryICML 2026
- Online Social Welfare Function-based Resource AllocationKanad Pardeshi, Samsara Foubert, Aarti SinghICML 2026
Builds on6
- Achieving Fairness in the Stochastic Multi-Armed Bandit ProblemVishakha Patil, Ganesh Ghalme, Vineet Nair, Y. NarahariAAAI 2020 · 131 citations
- Fair Algorithms for Multi-Agent Multi-Armed BanditsSafwan Hossain, Evi Micha, Nisarg ShahNeurIPS 2021 · 69 citations
- Universal and Tight Online Algorithms for Generalized-Mean WelfareSiddharth Barman, Arindam Khan, Arnab MaitiAAAI 2022 · 29 citations
- Fairness and Welfare Quantification for Regret in Multi-Armed BanditsSiddharth Barman, Arindam Khan, Arnab Maiti, Ayush SawarniAAAI 2023 · 18 citations
- Nash Regret Guarantees for Linear BanditsAyush Sawarni, Soumyabrata Pal, Siddharth BarmanNeurIPS 2023 · 12 citations
Related papers
- Honor Among Bandits: No-Regret Learning for Online Fair DivisionAriel D. Procaccia, Ben Schiffer, Shirley ZhangNeurIPS 2024 · 14 citations
- An Efficient Algorithm for Fair Multi-Agent Multi-Armed Bandit with Low RegretMatthew Jones, Huy L. Nguyen, Thy Dinh NguyenAAAI 2023 · 11 citations
- No-Regret Learning for Fair Multi-Agent Social Welfare OptimizationMengxiao Zhang, Ramiro Deo-Campo Vuong, Haipeng LuoNeurIPS 2024 · 7 citations
- Bandits with many optimal armsRianne de Heide, James Cheshire, Pierre Ménard, Alexandra CarpentierNeurIPS 2021 · 28 citations
- Player-optimal Stable Regret for Bandit Learning in Matching MarketsFang Kong, Shuai LiSODA 2023 · 6 citations
