Population-size-Aware Policy Optimization for Mean-Field Games
Pengdeng Li, Xinrun Wang, Shuxin Li, Hau Chan, Bo An
摘要
In this work, we attempt to bridge the two fields of finite-agent and infinite-agent games, by studying how the optimal policies of agents evolve with the number of agents (population size) in mean-field games, an agent-centric perspective in contrast to the existing works focusing typically on the convergence of the empirical distribution of the population. To this end, the premise is to obtain the optimal policies of a set of finite-agent games with different population sizes. However, either deriving the closed-form solution for each game is theoretically intractable, training a distinct policy for each game is computationally intensive, or directly applying the policy trained in a game to other games is sub-optimal. We address these challenges through the Population-size-Aware Policy Optimization (PAPO). Our contributions are three-fold. First, to efficiently generate efficient policies for games with different population sizes, we propose PAPO, which unifies two natural options (augmentation and hypernetwork) and achieves significantly better performance. PAPO consists of three components: i) the population-size encoding which transforms the original value of population size to an equivalent encoding to avoid training collapse, ii) a hypernetwork to generate a distinct policy for each game conditioned on the population size, and iii) the population size as an additional input to the generated policy. Next, we construct a multi-task-based training procedure to efficiently train the neural networks of PAPO by sampling data from multiple games with different population sizes. Finally, extensive experiments on multiple environments show the significant superiority of PAPO over baselines, and the analysis of the evolution of the generated policies further deepens our understanding of the two fields of finite-agent and infinite-agent games.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper21
- Provable Self-Play Algorithms for Competitive Reinforcement LearningYu Bai, Chi JinICML 2020 · 被引用 169 次
- Meta-Q-LearningRasool Fakoor, Pratik Chaudhari, Stefano Soatto, Alexander J. SmolaICLR 2020 · 被引用 162 次
- Global Convergence of Multi-Agent Policy Gradient in Markov Potential GamesStefanos Leonardos, Will Overman, Ioannis Panageas, Georgios PiliourasICLR 2022 · 被引用 158 次
- Fictitious Play for Mean Field Games: Continuous Time Analysis and ApplicationsSarah Perrin, Julien Pérolat, Mathieu Laurière, Matthieu Geist 等NeurIPS 2020 · 被引用 150 次
- Linear Last-iterate Convergence in Constrained Saddle-point OptimizationChen-Yu Wei, Chung-Wei Lee, Mengxiao Zhang, Haipeng LuoICLR 2021 · 被引用 146 次
相关 Paper
- Generalization in Mean Field Games by Learning Master PoliciesSarah Perrin, Mathieu Laurière, Julien Pérolat, Romuald Élie 等AAAI 2022 · 被引用 47 次
- Solving Continuous Mean Field Games: Deep Reinforcement Learning for Non-Stationary DynamicsLorenzo Magnino, Kai Shao, Zida Wu, Jiacheng Shen 等NeurIPS 2025 · 被引用 4 次
- Stochastic Semi-Gradient Descent for Learning Mean Field Games with Population-Aware Function ApproximationChenyu Zhang, Xu Chen, Xuan DiICLR 2025
- Population-Aware Imitation Learning in Mean-field Games with Common NoiseGrégoire Lambrecht, Mathieu LauriereICML 2026
- Coordinated Proximal Policy OptimizationZifan Wu, Chao Yu, Deheng Ye, Junge Zhang 等NeurIPS 2021 · 被引用 73 次
