Learning While Playing in Mean-Field Games: Convergence and Optimality
Qiaomin Xie, Zhuoran Yang, Zhaoran Wang, Andreea Minca
摘要
We study reinforcement learning in mean-field games. To achieve the Nash equilibrium, which consists of a policy and a mean-field state, existing algorithms require obtaining the optimal policy while fixing any mean-field state. In practice, however, the policy and the mean-field state evolve simultaneously, as each agent is learning while playing. To bridge such a gap, we propose a fictitious play algorithm, which alternatively updates the policy (learning) and the mean-field state (playing) by one step of policy optimization and gradient descent, respectively. Despite the nonstationarity induced by such an alternating scheme, we prove that the proposed algorithm converges to the Nash equilibrium with an explicit convergence rate. To the best of our knowledge, it is the first provably efficient algorithm that achieves learning while playing.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Scalable Deep Reinforcement Learning Algorithms for Mean Field GamesMathieu Laurière, Sarah Perrin, Sertan Girgin, Paul Muller 等ICML 2022 · 被引用 64 次
- Policy Mirror Ascent for Efficient and Independent Learning in Mean Field GamesBatuhan Yardim, Semih Cayci, Matthieu Geist, Niao HeICML 2023 · 被引用 33 次
- A Mean-Field Game Approach to Cloud Resource Management with Function ApproximationWeichao Mao, Haoran Qiu, Chen Wang, Hubertus Franke 等NeurIPS 2022 · 被引用 27 次
- Learning Regularized Monotone Graphon Mean-Field GamesFengzhuo Zhang, Vincent Y. F. Tan, Zhaoran Wang, Zhuoran YangNeurIPS 2023 · 被引用 14 次
- Learning Mean Field Games on Sparse Graphs: A Hybrid Graphex ApproachChristian Fabian, Kai Cui, Heinz KoepplICLR 2024 · 被引用 5 次
它引用的顶会 Paper5
- Adaptive Trust Region Policy Optimization: Global Convergence and Faster Rates for Regularized MDPsLior Shani, Yonathan Efroni, Shie MannorAAAI 2020 · 被引用 201 次
- Fictitious Play for Mean Field Games: Continuous Time Analysis and ApplicationsSarah Perrin, Julien Pérolat, Mathieu Laurière, Matthieu Geist 等NeurIPS 2020 · 被引用 150 次
- On the Convergence of Model Free Learning in Mean Field GamesRomuald Elie, Julien Pérolat, Mathieu Laurière, Matthieu Geist 等AAAI 2020 · 被引用 101 次
- Dynamic Regret of Policy Optimization in Non-Stationary EnvironmentsYingjie Fei, Zhuoran Yang, Zhaoran Wang, Qiaomin XieNeurIPS 2020 · 被引用 73 次
- Actor-Critic Provably Finds Nash Equilibria of Linear-Quadratic Mean-Field GamesZuyue Fu, Zhuoran Yang, Yongxin Chen, Zhaoran WangICLR 2020 · 被引用 61 次
相关 Paper
- Global Convergence of Policy Gradient for Linear-Quadratic Mean-Field Control/Game in Continuous TimeWeichen Wang, Jiequn Han, Zhuoran Yang, Zhaoran WangICML 2021 · 被引用 32 次
- Exponential Lower Bounds for Fictitious Play in Potential GamesIoannis Panageas, Nikolas Patris, Stratis Skoulakis, Volkan CevherNeurIPS 2023 · 被引用 1 次
- Last Iterate Convergence in Monotone Mean Field GamesNoboru Isobe, Kenshi Abe, Kaito AriuNeurIPS 2025 · 被引用 2 次
- Generalization in Mean Field Games by Learning Master PoliciesSarah Perrin, Mathieu Laurière, Julien Pérolat, Romuald Élie 等AAAI 2022 · 被引用 47 次
- Multi-Agent Meta-Reinforcement Learning: Sharper Convergence Rates with Task SimilarityWeichao Mao, Haoran Qiu, Chen Wang, Hubertus Franke 等NeurIPS 2023 · 被引用 17 次
