Policy Mirror Ascent for Efficient and Independent Learning in Mean Field Games
Batuhan Yardim, Semih Cayci, Matthieu Geist, Niao He
Abstract
Mean-field games have been used as a theoretical tool to obtain an approximate Nash equilibrium for symmetric and anonymous -player games. However, limiting applicability, existing theoretical results assume variations of a"population generative model", which allows arbitrary modifications of the population distribution by the learning algorithm. Moreover, learning algorithms typically work on abstract simulators with population instead of the -player game. Instead, we show that agents running policy mirror ascent converge to the Nash equilibrium of the regularized game within samples from a single sample trajectory without a population generative model, up to a standard error due to the mean field. Taking a divergent approach from the literature, instead of working with the best-response map we first show that a policy mirror ascent map can be used to construct a contractive operator having the Nash equilibrium as its fixed point. We analyze single-path TD learning for -agent games, proving sample complexity guarantees by only using a sample path from the -agent simulator without a population generative model. Furthermore, we demonstrate that our methodology allows for independent learning by agents with finite sample guarantees.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e0906b99-3531-4f2d-b625-822c55663779Cited by top-tier papers11
- A Novel Framework for Policy Mirror Descent with General Parameterization and Linear ConvergenceCarlo Alfano, Rui Yuan, Patrick RebeschiniNeurIPS 2023 · 25 citations
- Learning Regularized Monotone Graphon Mean-Field GamesFengzhuo Zhang, Vincent Y. F. Tan, Zhaoran Wang, Zhuoran YangNeurIPS 2023 · 14 citations
- On Imitation in Mean-field GamesGiorgia Ramponi, Pavel Kolev, Olivier Pietquin, Niao He et al.NeurIPS 2023 · 12 citations
- Major-Minor Mean Field Multi-Agent Reinforcement LearningKai Cui, Christian Fabian, Anam Tahir, Heinz KoepplICML 2024 · 6 citations
- Learning Discrete-Time Major-Minor Mean Field GamesKai Cui, Gökçe Dayanikli, Mathieu Laurière, Matthieu Geist et al.AAAI 2024 · 5 citations
Builds on11
- On the Global Convergence Rates of Softmax Policy Gradient MethodsJincheng Mei, Chenjun Xiao, Csaba Szepesvári, Dale SchuurmansICML 2020 · 349 citations
- Independent Policy Gradient Methods for Competitive Reinforcement LearningConstantinos Daskalakis, Dylan J. Foster, Noah GolowichNeurIPS 2020 · 200 citations
- Fictitious Play for Mean Field Games: Continuous Time Analysis and ApplicationsSarah Perrin, Julien Pérolat, Mathieu Laurière, Matthieu Geist et al.NeurIPS 2020 · 150 citations
- Mirror Descent Policy OptimizationManan Tomar, Lior Shani, Yonathan Efroni, Mohammad GhavamzadehICLR 2022 · 111 citations
- Decentralized Q-learning in Zero-sum Markov GamesMuhammed O. Sayin, Kaiqing Zhang, David S. Leslie, Tamer Basar et al.NeurIPS 2021 · 105 citations
Related papers
- On the Convergence of Model Free Learning in Mean Field GamesRomuald Elie, Julien Pérolat, Mathieu Laurière, Matthieu Geist et al.AAAI 2020 · 101 citations
- Actor-Critic Provably Finds Nash Equilibria of Linear-Quadratic Mean-Field GamesZuyue Fu, Zhuoran Yang, Yongxin Chen, Zhaoran WangICLR 2020 · 61 citations
- Learning in two-player zero-sum partially observable Markov games with perfect recallTadashi Kozuno, Pierre Ménard, Rémi Munos, Michal ValkoNeurIPS 2021 · 23 citations
- Multi-Agent Meta-Reinforcement Learning: Sharper Convergence Rates with Task SimilarityWeichao Mao, Haoran Qiu, Chen Wang, Hubertus Franke et al.NeurIPS 2023 · 17 citations
- Independent Policy Gradient for Large-Scale Markov Potential Games: Sharper Rates, Function Approximation, and Game-Agnostic ConvergenceDongsheng Ding, Chen-Yu Wei, Kaiqing Zhang, Mihailo R. JovanovicICML 2022 · 84 citations
