Exploration-Exploitation in Multi-Agent Competition: Convergence with Bounded Rationality
Stefanos Leonardos, Georgios Piliouras, Kelly Spendlove
摘要
The interplay between exploration and exploitation in competitive multi-agent learning is still far from being well understood. Motivated by this, we study smooth Q-learning, a prototypical learning model that explicitly captures the balance between game rewards and exploration costs. We show that Q-learning always converges to the unique quantal-response equilibrium (QRE), the standard solution concept for games under bounded rationality, in weighted zero-sum polymatrix games with heterogeneous learning agents using positive exploration rates. Complementing recent results about convergence in weighted potential games [14, 32] , we show that fast convergence of Q-learning in competitive settings is obtained regardless of the number of agents and without any need for parameter fine-tuning. As showcased by our experiments in network zero-sum games, these theoretical results provide the necessary guarantees for an algorithmic approach to the currently open problem of equilibrium selection in competitive multi-agent settings.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- Independent Policy Gradient for Large-Scale Markov Potential Games: Sharper Rates, Function Approximation, and Game-Agnostic ConvergenceDongsheng Ding, Chen-Yu Wei, Kaiqing Zhang, Mihailo R. JovanovicICML 2022 · 被引用 84 次
- Multi-Player Zero-Sum Markov Games with Networked Separable InteractionsChanwoo Park, Kaiqing Zhang, Asuman E. OzdaglarNeurIPS 2023 · 被引用 17 次
- Learning with Exposure Constraints in Recommendation SystemsOmer Ben-Porat, Rotem TorkanWWW 2023 · 被引用 16 次
- Approximating Nash Equilibria in Normal-Form Games via Stochastic OptimizationIan Gemp, Luke Marris, Georgios PiliourasICLR 2024 · 被引用 14 次
- Achieving Logarithmic Regret in KL-Regularized Zero-Sum Markov GamesAnupam Nayak, Tong Yang, Osman Yagan, Gauri Joshi 等ICML 2026 · 被引用 9 次
它引用的顶会 Paper6
- Independent Policy Gradient Methods for Competitive Reinforcement LearningConstantinos Daskalakis, Dylan J. Foster, Noah GolowichNeurIPS 2020 · 被引用 200 次
- Provable Self-Play Algorithms for Competitive Reinforcement LearningYu Bai, Chi JinICML 2020 · 被引用 169 次
- Linear Last-iterate Convergence in Constrained Saddle-point OptimizationChen-Yu Wei, Chung-Wei Lee, Mengxiao Zhang, Haipeng LuoICLR 2021 · 被引用 146 次
- From Poincaré Recurrence to Convergence in Imperfect Information Games: Finding Equilibrium via RegularizationJulien Pérolat, Rémi Munos, Jean-Baptiste Lespiau, Shayegan Omidshafiei 等ICML 2021 · 被引用 102 次
- Exploration-Exploitation in Multi-Agent Learning: Catastrophe Theory Meets Game TheoryStefanos Leonardos, Georgios PiliourasAAAI 2021 · 被引用 56 次
相关 Paper
- The Impact of Exploration on Convergence and Performance of Multi-Agent Q-Learning DynamicsAamal Abbas Hussain, Francesco Belardinelli, Dario PaccagnanICML 2023 · 被引用 2 次
- Stability of Multi-Agent Learning in Competitive Networks: Delaying the Onset of ChaosAamal Abbas Hussain, Francesco BelardinelliAAAI 2024 · 被引用 4 次
- Asymptotic Extinction in Large Coordination GamesDesmond Chan, Bart de Keijzer, Tobias Galla, Stefanos Leonardos 等AAAI 2025
- Asynchronous Gradient Play in Zero-Sum Multi-agent GamesRuicheng Ao, Shicong Cen, Yuejie ChiICLR 2023
- Provably Convergent Actor-Critic in Risk-averse MARLYizhou Zhang, Eric MazumdarICML 2026
