Competing for Shareable Arms in Multi-Player Multi-Armed Bandits
Renzhe Xu, Haotian Wang, Xingxuan Zhang, Bo Li, Peng Cui
摘要
Competitions for shareable and limited resources have long been studied with strategic agents. In reality, agents often have to learn and maximize the rewards of the resources at the same time. To design an individualized competing policy, we model the competition between agents in a novel multi-player multi-armed bandit (MPMAB) setting where players are selfish and aim to maximize their own rewards. In addition, when several players pull the same arm, we assume that these players averagely share the arms' rewards by expectation. Under this setting, we first analyze the Nash equilibrium when arms' rewards are known. Subsequently, we propose a novel Selfish MPMAB with Averaging Allocation (SMAA) approach based on the equilibrium. We theoretically demonstrate that SMAA could achieve a good regret guarantee for each player when all players follow the algorithm. Additionally, we establish that no single selfish player can significantly increase their rewards through deviation, nor can they detrimentally affect other players' rewards without incurring substantial losses for themselves. We finally validate the effectiveness of the method in extensive synthetic experiments.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- PPA-Game: Characterizing and Learning Competitive Dynamics Among Online Content CreatorsRenzhe Xu, Haotian Wang, Xingxuan Zhang, Bo Li 等KDD 2025 · 被引用 1 次
- Lower Bias, Higher Welfare: How Creator Competition Reshapes Bias-Variance Tradeoff in Recommendation Platforms?Kang Wang, Renzhe Xu, Bo LiKDD 2026
- Heterogeneous Data Game: Characterizing the Model Competition Across Multiple Data SourcesRenzhe Xu, Kang Wang, Bo LiICML 2025
- Multiple-play Stochastic Bandits with Prioritized Arm Capacity SharingHong Xie, Haoran Gu, Yanying Huang, Tao Tan 等AAAI 2026
它引用的顶会 Paper10
- Supply-Side Equilibria in Recommender SystemsMeena Jagadeesan, Nikhil Garg, Jacob SteinhardtNeurIPS 2023 · 被引用 53 次
- Learning Equilibria in Matching Markets from Bandit FeedbackMeena Jagadeesan, Alexander Wei, Yixin Wang, Michael I. Jordan 等NeurIPS 2021 · 被引用 52 次
- Beyond log2(T) regret for decentralized bandits in matching marketsSoumya Basu, Karthik Abinav Sankararaman, Abishek SankararamanICML 2021 · 被引用 45 次
- Heterogeneous Multi-player Multi-armed Bandits: Closing the Gap and GeneralizationChengshuai Shi, Wei Xiong, Cong Shen, Jing YangNeurIPS 2021 · 被引用 33 次
- Content Provider Dynamics and Coordination in Recommendation EcosystemsOmer Ben-Porat, Itay Rosenberg, Moshe TennenholtzNeurIPS 2020 · 被引用 22 次
相关 Paper
- Fair Algorithms for Multi-Agent Multi-Armed BanditsSafwan Hossain, Evi Micha, Nisarg ShahNeurIPS 2021 · 被引用 69 次
- Strategic Multi-Armed Bandit Problems Under Debt-Free ReportingAhmed Ben Yahmed, Clément Calauzènes, Vianney PerchetNeurIPS 2024 · 被引用 2 次
- Decentralized Scheduling with QoS Constraints: Achieving O(1) QoS Regret of Multi-Player BanditsQingsong Liu, Zhixuan FangAAAI 2024 · 被引用 5 次
- Optimal Algorithm for Max-Min Fair BanditZilong Wang, Zhiyao Zhang, Shuai LiICML 2025
- My Fair Bandit: Distributed Learning of Max-Min Fairness with Multi-player BanditsIlai Bistritz, Tavor Z. Baharav, Amir Leshem, Nicholas BambosICML 2020 · 被引用 40 次
