Policy Mirror Ascent for Efficient and Independent Learning in Mean Field Games
Batuhan Yardim, Semih Cayci, Matthieu Geist, Niao He
摘要
Mean-field games have been used as a theoretical tool to obtain an approximate Nash equilibrium for symmetric and anonymous -player games. However, limiting applicability, existing theoretical results assume variations of a"population generative model", which allows arbitrary modifications of the population distribution by the learning algorithm. Moreover, learning algorithms typically work on abstract simulators with population instead of the -player game. Instead, we show that agents running policy mirror ascent converge to the Nash equilibrium of the regularized game within samples from a single sample trajectory without a population generative model, up to a standard error due to the mean field. Taking a divergent approach from the literature, instead of working with the best-response map we first show that a policy mirror ascent map can be used to construct a contractive operator having the Nash equilibrium as its fixed point. We analyze single-path TD learning for -agent games, proving sample complexity guarantees by only using a sample path from the -agent simulator without a population generative model. Furthermore, we demonstrate that our methodology allows for independent learning by agents with finite sample guarantees.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- A Novel Framework for Policy Mirror Descent with General Parameterization and Linear ConvergenceCarlo Alfano, Rui Yuan, Patrick RebeschiniNeurIPS 2023 · 被引用 25 次
- Learning Regularized Monotone Graphon Mean-Field GamesFengzhuo Zhang, Vincent Y. F. Tan, Zhaoran Wang, Zhuoran YangNeurIPS 2023 · 被引用 14 次
- On Imitation in Mean-field GamesGiorgia Ramponi, Pavel Kolev, Olivier Pietquin, Niao He 等NeurIPS 2023 · 被引用 12 次
- Major-Minor Mean Field Multi-Agent Reinforcement LearningKai Cui, Christian Fabian, Anam Tahir, Heinz KoepplICML 2024 · 被引用 6 次
- Learning Discrete-Time Major-Minor Mean Field GamesKai Cui, Gökçe Dayanikli, Mathieu Laurière, Matthieu Geist 等AAAI 2024 · 被引用 5 次
它引用的顶会 Paper11
- On the Global Convergence Rates of Softmax Policy Gradient MethodsJincheng Mei, Chenjun Xiao, Csaba Szepesvári, Dale SchuurmansICML 2020 · 被引用 349 次
- Independent Policy Gradient Methods for Competitive Reinforcement LearningConstantinos Daskalakis, Dylan J. Foster, Noah GolowichNeurIPS 2020 · 被引用 200 次
- Fictitious Play for Mean Field Games: Continuous Time Analysis and ApplicationsSarah Perrin, Julien Pérolat, Mathieu Laurière, Matthieu Geist 等NeurIPS 2020 · 被引用 150 次
- Mirror Descent Policy OptimizationManan Tomar, Lior Shani, Yonathan Efroni, Mohammad GhavamzadehICLR 2022 · 被引用 111 次
- Decentralized Q-learning in Zero-sum Markov GamesMuhammed O. Sayin, Kaiqing Zhang, David S. Leslie, Tamer Basar 等NeurIPS 2021 · 被引用 105 次
相关 Paper
- On the Convergence of Model Free Learning in Mean Field GamesRomuald Elie, Julien Pérolat, Mathieu Laurière, Matthieu Geist 等AAAI 2020 · 被引用 101 次
- Actor-Critic Provably Finds Nash Equilibria of Linear-Quadratic Mean-Field GamesZuyue Fu, Zhuoran Yang, Yongxin Chen, Zhaoran WangICLR 2020 · 被引用 61 次
- Learning in two-player zero-sum partially observable Markov games with perfect recallTadashi Kozuno, Pierre Ménard, Rémi Munos, Michal ValkoNeurIPS 2021 · 被引用 23 次
- Multi-Agent Meta-Reinforcement Learning: Sharper Convergence Rates with Task SimilarityWeichao Mao, Haoran Qiu, Chen Wang, Hubertus Franke 等NeurIPS 2023 · 被引用 17 次
- Independent Policy Gradient for Large-Scale Markov Potential Games: Sharper Rates, Function Approximation, and Game-Agnostic ConvergenceDongsheng Ding, Chen-Yu Wei, Kaiqing Zhang, Mihailo R. JovanovicICML 2022 · 被引用 84 次
