Learning While Playing in Mean-Field Games: Convergence and Optimality
Qiaomin Xie, Zhuoran Yang, Zhaoran Wang, Andreea Minca
Abstract
We study reinforcement learning in mean-field games. To achieve the Nash equilibrium, which consists of a policy and a mean-field state, existing algorithms require obtaining the optimal policy while fixing any mean-field state. In practice, however, the policy and the mean-field state evolve simultaneously, as each agent is learning while playing. To bridge such a gap, we propose a fictitious play algorithm, which alternatively updates the policy (learning) and the mean-field state (playing) by one step of policy optimization and gradient descent, respectively. Despite the nonstationarity induced by such an alternating scheme, we prove that the proposed algorithm converges to the Nash equilibrium with an explicit convergence rate. To the best of our knowledge, it is the first provably efficient algorithm that achieves learning while playing.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cd7d55e6-e3d0-4fce-8143-58a0fbd46486Cited by top-tier papers9
- Scalable Deep Reinforcement Learning Algorithms for Mean Field GamesMathieu Laurière, Sarah Perrin, Sertan Girgin, Paul Muller et al.ICML 2022 · 64 citations
- Policy Mirror Ascent for Efficient and Independent Learning in Mean Field GamesBatuhan Yardim, Semih Cayci, Matthieu Geist, Niao HeICML 2023 · 33 citations
- A Mean-Field Game Approach to Cloud Resource Management with Function ApproximationWeichao Mao, Haoran Qiu, Chen Wang, Hubertus Franke et al.NeurIPS 2022 · 27 citations
- Learning Regularized Monotone Graphon Mean-Field GamesFengzhuo Zhang, Vincent Y. F. Tan, Zhaoran Wang, Zhuoran YangNeurIPS 2023 · 14 citations
- Learning Mean Field Games on Sparse Graphs: A Hybrid Graphex ApproachChristian Fabian, Kai Cui, Heinz KoepplICLR 2024 · 5 citations
Builds on5
- Adaptive Trust Region Policy Optimization: Global Convergence and Faster Rates for Regularized MDPsLior Shani, Yonathan Efroni, Shie MannorAAAI 2020 · 201 citations
- Fictitious Play for Mean Field Games: Continuous Time Analysis and ApplicationsSarah Perrin, Julien Pérolat, Mathieu Laurière, Matthieu Geist et al.NeurIPS 2020 · 150 citations
- On the Convergence of Model Free Learning in Mean Field GamesRomuald Elie, Julien Pérolat, Mathieu Laurière, Matthieu Geist et al.AAAI 2020 · 101 citations
- Dynamic Regret of Policy Optimization in Non-Stationary EnvironmentsYingjie Fei, Zhuoran Yang, Zhaoran Wang, Qiaomin XieNeurIPS 2020 · 73 citations
- Actor-Critic Provably Finds Nash Equilibria of Linear-Quadratic Mean-Field GamesZuyue Fu, Zhuoran Yang, Yongxin Chen, Zhaoran WangICLR 2020 · 61 citations
Related papers
- Global Convergence of Policy Gradient for Linear-Quadratic Mean-Field Control/Game in Continuous TimeWeichen Wang, Jiequn Han, Zhuoran Yang, Zhaoran WangICML 2021 · 32 citations
- Exponential Lower Bounds for Fictitious Play in Potential GamesIoannis Panageas, Nikolas Patris, Stratis Skoulakis, Volkan CevherNeurIPS 2023 · 1 citation
- Last Iterate Convergence in Monotone Mean Field GamesNoboru Isobe, Kenshi Abe, Kaito AriuNeurIPS 2025 · 2 citations
- Generalization in Mean Field Games by Learning Master PoliciesSarah Perrin, Mathieu Laurière, Julien Pérolat, Romuald Élie et al.AAAI 2022 · 47 citations
- Multi-Agent Meta-Reinforcement Learning: Sharper Convergence Rates with Task SimilarityWeichao Mao, Haoran Qiu, Chen Wang, Hubertus Franke et al.NeurIPS 2023 · 17 citations
