Optimistic Multi-Agent Policy Gradient
Wenshuai Zhao, Yi Zhao, Zhiyuan Li, Juho Kannala, Joni Pajarinen
Abstract
Relative overgeneralization (RO) occurs in cooperative multi-agent learning tasks when agents converge towards a suboptimal joint policy due to overfitting to suboptimal behavior of other agents. No methods have been proposed for addressing RO in multi-agent policy gradient (MAPG) methods although these methods produce state-of-the-art results. To address this gap, we propose a general, yet simple, framework to enable optimistic updates in MAPG methods that alleviate the RO problem. Our approach involves clipping the advantage to eliminate negative values, thereby facilitating optimistic updates in MAPG. The optimism prevents individual agents from quickly converging to a local optimum. Additionally, we provide a formal analysis to show that the proposed method retains optimality at a fixed point. In extensive evaluations on a diverse set of tasks including the Multi-agent MuJoCo and Overcooked benchmarks, our method outperforms strong baselines on 13 out of 19 tested tasks and matches the performance on the rest.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 05d0d6bb-b8f3-4132-b023-5317e11b64bcCited by top-tier papers2
- AgentMixer: Multi-Agent Correlated Policy FactorizationZhiyuan Li, Wenshuai Zhao, Lijun Wu, Joni PajarinenAAAI 2025 · 7 citations
- Learning Progress Driven Multi-Agent CurriculumWenshuai Zhao, Zhiyuan Li, Joni PajarinenICML 2025
Builds on6
- Weighted QMIX: Expanding Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement LearningTabish Rashid, Gregory Farquhar, Bei Peng, Shimon WhitesonNeurIPS 2020 · 1,960 citations
- Cooperative Exploration for Multi-Agent Deep Reinforcement LearningIou-Jen Liu, Unnat Jain, Raymond A. Yeh, Alexander G. SchwingICML 2021 · 133 citations
- PMIC: Improving Multi-Agent Reinforcement Learning with Progressive Mutual Information CollaborationPengyi Li, Hongyao Tang, Tianpei Yang, Xiaotian Hao et al.ICML 2022 · 49 citations
- An operator view of policy gradient methodsDibya Ghosh, Marlos C. Machado, Nicolas Le RouxNeurIPS 2020 · 30 citations
- Learning Zero-Shot Cooperation with Humans, Assuming Humans Are BiasedChao Yu, Jiaxuan Gao, Weilin Liu, Botian Xu et al.ICLR 2023 · 4 citations
Related papers
- Trust Region Policy Optimisation in Multi-Agent Reinforcement LearningJakub Grudzien Kuba, Ruiqing Chen, Muning Wen, Ying Wen et al.ICLR 2022 · 367 citations
- Local Optimization Achieves Global Optimality in Multi-Agent Reinforcement LearningYulai Zhao, Zhuoran Yang, Zhaoran Wang, Jason D. LeeICML 2023 · 8 citations
- Optimistic Value Instructors for Cooperative Multi-Agent Reinforcement LearningChao Li, Yupeng Zhang, Jianqi Wang, Yujing Hu et al.AAAI 2024 · 4 citations
- Robust and Diverse Multi-Agent Learning via Rational Policy GradientNiklas Lauffer, Ameesh Shah, Micah Carroll, Sanjit A. Seshia et al.NeurIPS 2025 · 4 citations
- Learning Explicit Credit Assignment for Cooperative Multi-Agent Reinforcement Learning via Polarization Policy GradientWubing Chen, Wenbin Li, Xiao Liu, Shangdong Yang et al.AAAI 2023 · 11 citations
