On the convergence of policy gradient methods to Nash equilibria in general stochastic games
Angeliki Giannou, Kyriakos Lotidis, Panayotis Mertikopoulos, Emmanouil V. Vlatakis-Gkaragkounis
摘要
Learning in stochastic games is a notoriously difficult problem because, in addition to each other's strategic decisions, the players must also contend with the fact that the game itself evolves over time, possibly in a very complicated manner. Because of this, the convergence properties of popular learning algorithms - like policy gradient and its variants - are poorly understood, except in specific classes of games (such as potential or two-player, zero-sum games). In view of this, we examine the long-run behavior of policy gradient methods with respect to Nash equilibrium policies that are second-order stationary (SOS) in a sense similar to the type of sufficiency conditions used in optimization. Our first result is that SOS policies are locally attracting with high probability, and we show that policy gradient trajectories with gradient estimates provided by the REINFORCE algorithm achieve an distance-squared convergence rate if the method's step-size is chosen appropriately. Subsequently, specializing to the class of deterministic Nash policies, we show that this rate can be improved dramatically and, in fact, policy gradient methods converge within a finite number of iterations in that case.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Optimistic Policy Gradient in Multi-Player Markov Games with a Single Controller: Convergence beyond the Minty PropertyIoannis Anagnostides, Ioannis Panageas, Gabriele Farina, Tuomas SandholmAAAI 2024 · 被引用 3 次
- Policy Gradient Methods Converge Globally in Imperfect-Information Extensive-Form GamesFivos Kalogiannis, Gabriele FarinaNeurIPS 2025 · 被引用 2 次
- Robust Equilibria in Continuous Games: From Strategic to Dynamic RobustnessKyriakos Lotidis, Panayotis Mertikopoulos, Nicholas Bambos, Jose H. BlanchetNeurIPS 2025 · 被引用 2 次
- Shuffling the Data, Extrapolating the Step: Sharper Bias In Constant Step-Size SGDKonstantinos Emmanouilidis, Emmanouil-Vasileios Vlatakis-Gkaragkounis, Rene VidalICLR 2026
- Learning to Steer Markovian Agents under Model UncertaintyJiawei Huang, Vinzenz Thoma, Zebang Shen, Heinrich H. Nax 等ICLR 2025
它引用的顶会 Paper8
- Independent Policy Gradient Methods for Competitive Reinforcement LearningConstantinos Daskalakis, Dylan J. Foster, Noah GolowichNeurIPS 2020 · 被引用 200 次
- Global Convergence of Multi-Agent Policy Gradient in Markov Potential GamesStefanos Leonardos, Will Overman, Ioannis Panageas, Georgios PiliourasICLR 2022 · 被引用 158 次
- Decentralized Q-learning in Zero-sum Markov GamesMuhammed O. Sayin, Kaiqing Zhang, David S. Leslie, Tamer Basar 等NeurIPS 2021 · 被引用 105 次
- Fast Policy Extragradient Methods for Competitive Games with Entropy RegularizationShicong Cen, Yuting Wei, Yuejie ChiNeurIPS 2021 · 被引用 105 次
- No-Regret Learning and Mixed Nash Equilibria: They Do Not MixEmmanouil V. Vlatakis-Gkaragkounis, Lampros Flokas, Thanasis Lianeas, Panayotis Mertikopoulos 等NeurIPS 2020 · 被引用 100 次
相关 Paper
- Independent Policy Gradient for Large-Scale Markov Potential Games: Sharper Rates, Function Approximation, and Game-Agnostic ConvergenceDongsheng Ding, Chen-Yu Wei, Kaiqing Zhang, Mihailo R. JovanovicICML 2022 · 被引用 84 次
- A Policy-Gradient Approach to Solving Imperfect-Information Games with Best-Iterate ConvergenceMingyang Liu, Gabriele Farina, Asuman E. OzdaglarICLR 2025
- Convergence and Optimality of Policy Gradient Methods in Weakly Smooth SettingsMatthew Shunshi Zhang, Murat A. Erdogdu, Animesh GargAAAI 2022 · 被引用 6 次
- Policy Optimization for Markov Games: Unified Framework and Faster ConvergenceRunyu Zhang, Qinghua Liu, Huan Wang, Caiming Xiong 等NeurIPS 2022 · 被引用 32 次
- Global Convergence of Policy Gradient for Linear-Quadratic Mean-Field Control/Game in Continuous TimeWeichen Wang, Jiequn Han, Zhuoran Yang, Zhaoran WangICML 2021 · 被引用 32 次
