Major-Minor Mean Field Multi-Agent Reinforcement Learning
Kai Cui, Christian Fabian, Anam Tahir, Heinz Koeppl
Abstract
Multi-agent reinforcement learning (MARL) remains difficult to scale to many agents. Recent MARL using Mean Field Control (MFC) provides a tractable and rigorous approach to otherwise difficult cooperative MARL. However, the strict MFC assumption of many independent, weakly-interacting agents is too inflexible in practice. We generalize MFC to instead simultaneously model many similar and few complex agents -- as Major-Minor Mean Field Control (M3FC). Theoretically, we give approximation results for finite agent control, and verify the sufficiency of stationary policies for optimality together with a dynamic programming principle. Algorithmically, we propose Major-Minor Mean Field MARL (M3FMARL) for finite agent systems instead of the limiting system. The algorithm is shown to approximate the policy gradient of the underlying M3FC MDP. Finally, we demonstrate its capabilities experimentally in various scenarios. We observe a strong performance in comparison to state-of-the-art policy gradient MARL methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on9
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Fictitious Play for Mean Field Games: Continuous Time Analysis and ApplicationsSarah Perrin, Julien Pérolat, Mathieu Laurière, Matthieu Geist et al.NeurIPS 2020 · 150 citations
- Scalable Multi-Agent Reinforcement Learning for Networked Systems with Average RewardGuannan Qu, Yiheng Lin, Adam Wierman, Na LiNeurIPS 2020 · 99 citations
- Learning Graphon Mean Field Games and Approximate Nash EquilibriaKai Cui, Heinz KoepplICLR 2022 · 50 citations
- Generalization in Mean Field Games by Learning Master PoliciesSarah Perrin, Mathieu Laurière, Julien Pérolat, Romuald Élie et al.AAAI 2022 · 47 citations
Related papers
- Global Convergence of Policy Gradient for Linear-Quadratic Mean-Field Control/Game in Continuous TimeWeichen Wang, Jiequn Han, Zhuoran Yang, Zhaoran WangICML 2021 · 32 citations
- Learning Decentralized Partially Observable Mean Field Control for Artificial Collective BehaviorKai Cui, Sascha Hauck, Christian Fabian, Heinz KoepplICLR 2024 · 13 citations
- Breaking the Curse of Many Agents: Provable Mean Embedding Q-Iteration for Mean-Field Reinforcement LearningLingxiao Wang, Zhuoran Yang, Zhaoran WangICML 2020 · 31 citations
- Decentralized Mean Field GamesSriram Ganapathi Subramanian, Matthew E. Taylor, Mark Crowley, Pascal PoupartAAAI 2022 · 19 citations
- Actor-Critic Provably Finds Nash Equilibria of Linear-Quadratic Mean-Field GamesZuyue Fu, Zhuoran Yang, Yongxin Chen, Zhaoran WangICLR 2020 · 61 citations
