Paths to Equilibrium in Games
Bora Yongacoglu, Gürdal Arslan, Lacra Pavel, Serdar Yüksel
摘要
In multi-agent reinforcement learning (MARL) and game theory, agents repeatedly interact and revise their strategies as new data arrives, producing a sequence of strategy profiles. This paper studies sequences of strategies satisfying a pairwise constraint inspired by policy updating in reinforcement learning, where an agent who is best responding in one period does not switch its strategy in the next period. This constraint merely requires that optimizing agents do not switch strategies, but does not constrain the non-optimizing agents in any way, and thus allows for exploration. Sequences with this property are called satisficing paths, and arise naturally in many MARL algorithms. A fundamental question about strategic dynamics is such: for a given game and initial strategy profile, is it always possible to construct a satisficing path that terminates at an equilibrium? The resolution of this question has implications about the capabilities or limitations of a class of MARL algorithms. We answer this question in the affirmative for normal-form games. Our analysis reveals a counterintuitive insight that reward deteriorating strategic updates are key to driving play to equilibrium along a satisficing path.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Tree-Based Stochastic Optimization for Solving Large-Scale Urban Network Security GamesShuxin Zhuang, Linjian Meng, Shuxin Li, Minming Li 等AAAI 2026
- Reducing Variance of Stochastic Optimization for Approximating Nash Equilibria in Normal-Form GamesLinjian Meng, Wubing Chen, Wenbin Li, Tianpei Yang 等ICML 2025
它引用的顶会 Paper5
- Independent Policy Gradient Methods for Competitive Reinforcement LearningConstantinos Daskalakis, Dylan J. Foster, Noah GolowichNeurIPS 2020 · 被引用 200 次
- No-Regret Learning and Mixed Nash Equilibria: They Do Not MixEmmanouil V. Vlatakis-Gkaragkounis, Lampros Flokas, Thanasis Lianeas, Panayotis Mertikopoulos 等NeurIPS 2020 · 被引用 100 次
- Adversarial Example GamesAvishek Joey Bose, Gauthier Gidel, Hugo Berard, Andre Cianflone 等NeurIPS 2020 · 被引用 59 次
- Two-Scale Gradient Descent Ascent Dynamics Finds Mixed Nash Equilibria of Continuous Games: A Mean-Field PerspectiveYulong LuICML 2023 · 被引用 31 次
- Fictitious Play and Best-Response Dynamics in Identical Interest and Zero-Sum Stochastic GamesLucas Baudin, Rida LarakiICML 2022 · 被引用 20 次
相关 Paper
- Learning While Playing in Mean-Field Games: Convergence and OptimalityQiaomin Xie, Zhuoran Yang, Zhaoran Wang, Andreea MincaICML 2021 · 被引用 45 次
- Exponential Lower Bounds for Fictitious Play in Potential GamesIoannis Panageas, Nikolas Patris, Stratis Skoulakis, Volkan CevherNeurIPS 2023 · 被引用 1 次
- Fast computation of Nash Equilibria in Imperfect Information GamesRémi Munos, Julien Pérolat, Jean-Baptiste Lespiau, Mark Rowland 等ICML 2020 · 被引用 11 次
- Is Learning in Games Good for the Learners?William Brown, Jon Schneider, Kiran VodrahalliNeurIPS 2023 · 被引用 27 次
- The Impact of Exploration on Convergence and Performance of Multi-Agent Q-Learning DynamicsAamal Abbas Hussain, Francesco Belardinelli, Dario PaccagnanICML 2023 · 被引用 2 次
