Cooperative Multi-player Bandit Optimization
Ilai Bistritz, Nicholas Bambos
摘要
Consider a team of cooperative players that take actions in a networkedenvironment. At each turn, each player chooses an action and receives a reward that is an unknown function of all the players' actions. The goal of the team of players is to learn to play together the action profile that maximizes the sum of their rewards. However, players cannot observe the actions or rewards of other players, and can only get this information by communicating with their neighbors. We design a distributed learning algorithm that overcomes the informational bias players have towards maximizing the rewards of nearby players they got more information about. We assume twice continuously differentiable reward functions and constrained convex and compact action sets. Our communication graph is a random time-varying graph that follows an ergodic Markov chain. We prove that even if at every turn players take actions based only on the small random subset of the players' rewards that they know, our algorithm converges with probability 1 to the set of stationary points of (projected) gradient ascent on the sum of rewards function. Hence, if the sum of rewards is concave, then the algorithm converges with probability 1 to an optimal action profile.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Global Convergence of Multi-Agent Policy Gradient in Markov Potential GamesStefanos Leonardos, Will Overman, Ioannis Panageas, Georgios PiliourasICLR 2022 · 被引用 158 次
- Independent Policy Gradient for Large-Scale Markov Potential Games: Sharper Rates, Function Approximation, and Game-Agnostic ConvergenceDongsheng Ding, Chen-Yu Wei, Kaiqing Zhang, Mihailo R. JovanovicICML 2022 · 被引用 84 次
- Cooperative Stochastic Bandits with Asynchronous Agents and Constrained FeedbackLin Yang, Yu-Zhen Janice Chen, Stephen Pasteris, Mohammad H. Hajiesmaili 等NeurIPS 2021 · 被引用 36 次
- Online Learning for Load Balancing of Unknown Monotone Resource Allocation GamesIlai Bistritz, Nicholas BambosICML 2021 · 被引用 9 次
- Micro-MAMA: Multi-Agent Reinforcement Learning for Multicore PrefetchingCharles Block, Gerasimos Gerogiannis, Josep TorrellasMICRO 2025 · 被引用 6 次
它引用的顶会 Paper1
相关 Paper
- Queue Up Your Regrets: Achieving the Dynamic Capacity Region of Multiplayer BanditsIlai Bistritz, Nicholas BambosNeurIPS 2022 · 被引用 5 次
- Decentralized Policy Gradient Descent Ascent for Safe Multi-Agent Reinforcement LearningSongtao Lu, Kaiqing Zhang, Tianyi Chen, Tamer Basar 等AAAI 2021 · 被引用 93 次
- Interactive Inverse Reinforcement Learning for Cooperative GamesThomas Kleine Büning, Anne-Marie George, Christos DimitrakakisICML 2022 · 被引用 7 次
- Optimally Improving Cooperative Learning in a Social SettingShahrzad Haddadan, Cheng Xin, Jie GaoICML 2024 · 被引用 2 次
- Online Convex Optimization Over Erdos-Renyi Random NetworksJinlong Lei, Peng Yi, Yiguang Hong, Jie Chen 等NeurIPS 2020 · 被引用 25 次
