Exploration-Exploitation in Multi-Agent Learning: Catastrophe Theory Meets Game Theory
Stefanos Leonardos, Georgios Piliouras
摘要
Exploration-exploitation is a powerful and practical tool in multi-agent learning (MAL), however, its effects are far from understood. To make progress in this direction, we study a smooth analogue of Q-learning. We start by showing that our learning model has strong theoretical justification as an optimal model for studying exploration-exploitation. Specifically, we prove that smooth Q-learning has bounded regret in arbitrary games for a cost model that explicitly captures the balance between game and exploration costs and that it always converges to the set of quantal-response equilibria (QRE), the standard solution concept for games under bounded rationality, in weighted potential games with heterogeneous learning agents. In our main task, we then turn to measure the effect of exploration in collective system performance. We characterize the geometry of the QRE surface in low-dimensional MAL systems and link our findings with catastrophe (bifurcation) theory. In particular, as the exploration hyperparameter evolves over-time, the system undergoes phase transitions where the number and stability of equilibria can change radically given an infinitesimal change to the exploration parameter. Based on this, we provide a formal theoretical treatment of how tuning the exploration parameter can provably lead to equilibrium selection with both positive as well as negative (and potentially unbounded) effects to system performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Independent Policy Gradient for Large-Scale Markov Potential Games: Sharper Rates, Function Approximation, and Game-Agnostic ConvergenceDongsheng Ding, Chen-Yu Wei, Kaiqing Zhang, Mihailo R. JovanovicICML 2022 · 被引用 84 次
- Exploration-Exploitation in Multi-Agent Competition: Convergence with Bounded RationalityStefanos Leonardos, Georgios Piliouras, Kelly SpendloveNeurIPS 2021 · 被引用 43 次
- The Best of Both Worlds in Network Population Games: Reaching Consensus and Convergence to EquilibriumShuyue Hu, Harold Soh, Georgios PiliourasNeurIPS 2023 · 被引用 9 次
- Beating Price of Anarchy and Gradient Descent without Regret in Potential GamesIosif Sakos, Stefanos Leonardos, Stelios Andrew Stavroulakis, Will Overman 等ICLR 2024 · 被引用 3 次
- The Impact of Exploration on Convergence and Performance of Multi-Agent Q-Learning DynamicsAamal Abbas Hussain, Francesco Belardinelli, Dario PaccagnanICML 2023 · 被引用 2 次
它引用的顶会 Paper3
- Provable Self-Play Algorithms for Competitive Reinforcement LearningYu Bai, Chi JinICML 2020 · 被引用 169 次
- From Poincaré Recurrence to Convergence in Imperfect Information Games: Finding Equilibrium via RegularizationJulien Pérolat, Rémi Munos, Jean-Baptiste Lespiau, Shayegan Omidshafiei 等ICML 2021 · 被引用 102 次
- Smooth markets: A basic mechanism for organizing gradient-based learnersDavid Balduzzi, Wojciech M. Czarnecki, Tom Anthony, Ian Gemp 等ICLR 2020 · 被引用 15 次
相关 Paper
- Asymptotic Extinction in Large Coordination GamesDesmond Chan, Bart de Keijzer, Tobias Galla, Stefanos Leonardos 等AAAI 2025
- Tractable Multi-Agent Reinforcement Learning through Behavioral EconomicsEric Mazumdar, Kishan Panaganti, Laixi ShiICLR 2025
- Stability of Multi-Agent Learning in Competitive Networks: Delaying the Onset of ChaosAamal Abbas Hussain, Francesco BelardinelliAAAI 2024 · 被引用 4 次
- Robust Adversarial Reinforcement Learning via Bounded Rationality CurriculaAryaman Reddi, Maximilian Tölle, Jan Peters, Georgia Chalvatzaki 等ICLR 2024 · 被引用 11 次
- Provably Convergent Actor-Critic in Risk-averse MARLYizhou Zhang, Eric MazumdarICML 2026
