Rethinking Individual Global Max in Cooperative Multi-Agent Reinforcement Learning
Yitian Hong, Yaochu Jin, Yang Tang
Abstract
In cooperative multi-agent reinforcement learning, centralized training and decentralized execution (CTDE) has achieved remarkable success. Individual Global Max (IGM) decomposition, which is an important element of CTDE, measures the consistency between local and joint policies. The majority of IGM-based research focuses on how to establish this consistent relationship, but little attention has been paid to examining IGM's potential flaws. In this work, we reveal that the IGM condition is a lossy decomposition, and the error of lossy decomposition will accumulated in hypernetwork-based methods. To address the above issue, we propose to adopt an imitation learning strategy to separate the lossy decomposition from Bellman iterations, thereby avoiding error accumulation. The proposed strategy is theoretically proved and empirically verified on the StarCraft Multi-Agent Challenge benchmark problem with zero sight view. The results also confirm that the proposed method outperforms state-of-the-art IGM-based approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3048cdf5-e9e4-4294-ba5c-8d1d296ffbb6Cited by top-tier papers6
- Understanding Individual Agent Importance in Multi-Agent System via Counterfactual ReasoningJianming Chen, Yawen Wang, Junjie Wang, Xiaofei Xie et al.AAAI 2025 · 11 citations
- AgentMixer: Multi-Agent Correlated Policy FactorizationZhiyuan Li, Wenshuai Zhao, Lijun Wu, Joni PajarinenAAAI 2025 · 7 citations
- Multi-agent Markov EntanglementShuze Chen, Tianyi PengNeurIPS 2025
- HMARL-CBF - Hierarchical Multi-Agent Reinforcement Learning with Control Barrier Functions for Safety-Critical Autonomous SystemsH. M. Sabbir Ahmad, Ehsan Sabouni, Alexander Wasilkoff, Param Budhraja et al.NeurIPS 2025
- LC-ERD: Mining Latent Logic for Self-Evolving Reasoning via Consistency-Regulated Reward DecompositionYanyu Chen, Jiyue Jiang, Dianzhi Yu, Zheng Wu et al.KDD 2026
Builds on7
- QPLEX: Duplex Dueling Multi-Agent Q-LearningJianhao Wang, Zhizhou Ren, Terry Liu, Yang Yu et al.ICLR 2021 · 595 citations
- ROMA: Multi-Agent Reinforcement Learning with Emergent RolesTonghan Wang, Heng Dong, Victor R. Lesser, Chongjie ZhangICML 2020 · 286 citations
- Multi-Agent Reinforcement Learning for Active Voltage Control on Power Distribution NetworksJianhong Wang, Wangkun Xu, Yunjie Gu, Wenbin Song et al.NeurIPS 2021 · 216 citations
- Learning Nearly Decomposable Value Functions Via Communication MinimizationTonghan Wang, Jianhao Wang, Chongyi Zheng, Chongjie ZhangICLR 2020 · 170 citations
- DFAC Framework: Factorizing the Value Function via Quantile Mixture for Multi-Agent Distributional Q-LearningWei-Fang Sun, Cheng-Kuang Lee, Chun-Yi LeeICML 2021 · 56 citations
Related papers
- Beyond Monotonicity: Revisiting Factorization Principles in Multi-Agent Q-LearningTianmeng Hu, Yongzheng Cui, Rui Tang, Biao Luo et al.AAAI 2026
- Revisiting Some Common Practices in Cooperative Multi-Agent Reinforcement LearningWei Fu, Chao Yu, Zelai Xu, Jiaqi Yang et al.ICML 2022 · 49 citations
- Dual Self-Awareness Value Decomposition Framework without Individual Global Max for Cooperative MARLZhiwei Xu, Bin Zhang, Dapeng Li, Guangchong Zhou et al.NeurIPS 2023 · 12 citations
- Learning Implicit Credit Assignment for Cooperative Multi-Agent Reinforcement LearningMeng Zhou, Ziyu Liu, Pengwei Sui, Yixuan Li et al.NeurIPS 2020 · 142 citations
- Multiagent Q-learning with Sub-Team CoordinationWenhan Huang, Kai Li, Kun Shao, Tianze Zhou et al.NeurIPS 2022 · 12 citations
