Greedy based Value Representation for Optimal Coordination in Multi-agent Reinforcement Learning
Lipeng Wan, Zeyang Liu, Xingyu Chen, Xuguang Lan, Nanning Zheng
Abstract
Due to the representation limitation of the joint Q value function, multi-agent reinforcement learning methods with linear value decomposition (LVD) or monotonic value decomposition (MVD) suffer from relative overgeneralization. As a result, they can not ensure optimal consistency (i.e., the correspondence between individual greedy actions and the maximal true Q value). In this paper, we derive the expression of the joint Q value function of LVD and MVD. According to the expression, we draw a transition diagram, where each self-transition node (STN) is a possible convergence. To ensure optimal consistency, the optimal node is required to be the unique STN. Therefore, we propose the greedy-based value representation (GVR), which turns the optimal node into an STN via inferior target shaping and further eliminates the non-optimal STNs via superior experience replay. In addition, GVR achieves an adaptive trade-off between optimality and stability. Our method outperforms state-of-the-art baselines in experiments on various benchmarks. Theoretical proofs and empirical results on matrix games demonstrate that GVR ensures optimal consistency under sufficient exploration.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 16d0bf3d-9a5b-4d41-b20a-a4b2148b74fcCited by top-tier papers6
- Counterfactual Conservative Q Learning for Offline Multi-agent Reinforcement LearningJianzhun Shao, Yun Qu, Chen Chen, Hongchang Zhang et al.NeurIPS 2023 · 56 citations
- Automatic Grouping for Efficient Cooperative Multi-Agent Reinforcement LearningYifan Zang, Jinmin He, Kai Li, Haobo Fu et al.NeurIPS 2023 · 37 citations
- Complementary Attention for Multi-Agent Reinforcement LearningJianzhun Shao, Hongchang Zhang, Yun Qu, Chang Liu et al.ICML 2023 · 17 citations
- Retaining Suboptimal Actions to Follow Shifting Optima in Multi-Agent Reinforcement LearningYonghyeon Jo, Sunwoo Lee, Seungyul HanICLR 2026 · 5 citations
- Backpropagation Through AgentsZhiyuan Li, Wenshuai Zhao, Lijun Wu, Joni PajarinenAAAI 2024 · 3 citations
Builds on4
- Weighted QMIX: Expanding Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement LearningTabish Rashid, Gregory Farquhar, Bei Peng, Shimon WhitesonNeurIPS 2020 · 1,960 citations
- Deep Coordination GraphsWendelin Boehmer, Vitaly Kurin, Shimon WhitesonICML 2020 · 209 citations
- Q-value Path Decomposition for Deep Multiagent Reinforcement LearningYaodong Yang, Jianye Hao, Guangyong Chen, Hongyao Tang et al.ICML 2020 · 64 citations
- UneVEn: Universal Value Exploration for Multi-Agent Reinforcement LearningTarun Gupta, Anuj Mahajan, Bei Peng, Wendelin Boehmer et al.ICML 2021 · 59 citations
Related papers
- Beyond Monotonicity: Revisiting Factorization Principles in Multi-Agent Q-LearningTianmeng Hu, Yongzheng Cui, Rui Tang, Biao Luo et al.AAAI 2026
- Non-Linear Coordination GraphsYipeng Kang, Tonghan Wang, Qianlan Yang, Xiaoran Wu et al.NeurIPS 2022 · 14 citations
- ResQ: A Residual Q Function-based Approach for Multi-Agent Reinforcement Learning Value FactorizationSiqi Shen, Mengwei Qiu, Jun Liu, Weiquan Liu et al.NeurIPS 2022 · 35 citations
- Variational Empowerment as Representation Learning for Goal-Conditioned Reinforcement LearningJongwook Choi, Archit Sharma, Honglak Lee, Sergey Levine et al.ICML 2021 · 41 citations
- Potentially Optimal Joint Actions Recognition for Cooperative Multi-Agent Reinforcement LearningChang Huang, Shatong Zhu, Junqiao Zhao, Hongtu Zhou et al.ICLR 2026
