Gradient-Protected Value Decomposition for Cooperative Multi-Agent Reinforcement Learning
Jie Hou, Haowen Dou, Lujuan Dang, Liangjun Chen, Chenyang Ge
Abstract
In recent years, deep multi-agent reinforcement learning (MARL) has demonstrated remarkable potential in solving complex cooperative tasks by enabling decentralized yet efficient coordination among agents. However, during decentralized training, agent policy updates induced by different joint action samples may conflict, leading to gradient interference that hinders convergence and the emergence of coordinated behavior. In this paper, we analyze and empirically validate the phenomenon of gradient interference. To address this, we then propose Gradient-Protected Value Decomposition (GPVD), a novel MARL framework that explicitly protects the gradient signals of optimal collaborative actions by suppressing the impact of interfering actions. GPVD employs a dynamic gradient protection mechanism that identifies optimal collaborative joint actions and reweights the loss to attenuate gradients from non-collaborative interfering actions. To effectively identify high-value collaborative actions, we apply SimHash-based state grouping to discover consistent collaboration patterns across similar states. Furthermore, a countbased intrinsic reward is incorporated to encourage exploration and improve the coverage of potentially optimal joint actions. Experiments on challenging multi-agent benchmarks demonstrate that GPVD achieves faster convergence, stronger coordination, and greater training stability compared to stateof-the-art value decomposition methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c3b1054c-a3e5-46d6-a981-cf3596789f79Builds on13
- Weighted QMIX: Expanding Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement LearningTabish Rashid, Gregory Farquhar, Bei Peng, Shimon WhitesonNeurIPS 2020 · 1,960 citations
- Multi-Agent Reinforcement Learning for Active Voltage Control on Power Distribution NetworksJianhong Wang, Wangkun Xu, Yunjie Gu, Wenbin Song et al.NeurIPS 2021 · 216 citations
- Episodic Multi-agent Reinforcement Learning with Curiosity-driven ExplorationLulu Zheng, Jiarui Chen, Jianhao Wang, Jiamin He et al.NeurIPS 2021 · 126 citations
- MASER: Multi-Agent Reinforcement Learning with Subgoals Generated from Experience Replay BufferJeewon Jeon, Woojun Kim, Whiyoung Jung, Youngchul SungICML 2022 · 53 citations
- ResQ: A Residual Q Function-based Approach for Multi-Agent Reinforcement Learning Value FactorizationSiqi Shen, Mengwei Qiu, Jun Liu, Weiquan Liu et al.NeurIPS 2022 · 35 citations
Related papers
- Q-value Path Decomposition for Deep Multiagent Reinforcement LearningYaodong Yang, Jianye Hao, Guangyong Chen, Hongyao Tang et al.ICML 2020 · 64 citations
- Revisiting Some Common Practices in Cooperative Multi-Agent Reinforcement LearningWei Fu, Chao Yu, Zelai Xu, Jiaqi Yang et al.ICML 2022 · 49 citations
- HAVEN: Hierarchical Cooperative Multi-Agent Reinforcement Learning with Dual Coordination MechanismZhiwei Xu, Yunpeng Bai, Bin Zhang, Dapeng Li et al.AAAI 2023 · 46 citations
- Evolutionary Population Curriculum for Scaling Multi-Agent Reinforcement LearningQian Long, Zihan Zhou, Abhinav Gupta, Fei Fang et al.ICLR 2020
- Conditional Diffusion Model for Multi-Agent Dynamic Task DecompositionYanda Zhu, Yuanyang Zhu, Daoyi Dong, Caihua Chen et al.AAAI 2026
