Greedy based Value Representation for Optimal Coordination in Multi-agent Reinforcement Learning
Lipeng Wan, Zeyang Liu, Xingyu Chen, Xuguang Lan, Nanning Zheng
摘要
Due to the representation limitation of the joint Q value function, multi-agent reinforcement learning methods with linear value decomposition (LVD) or monotonic value decomposition (MVD) suffer from relative overgeneralization. As a result, they can not ensure optimal consistency (i.e., the correspondence between individual greedy actions and the maximal true Q value). In this paper, we derive the expression of the joint Q value function of LVD and MVD. According to the expression, we draw a transition diagram, where each self-transition node (STN) is a possible convergence. To ensure optimal consistency, the optimal node is required to be the unique STN. Therefore, we propose the greedy-based value representation (GVR), which turns the optimal node into an STN via inferior target shaping and further eliminates the non-optimal STNs via superior experience replay. In addition, GVR achieves an adaptive trade-off between optimality and stability. Our method outperforms state-of-the-art baselines in experiments on various benchmarks. Theoretical proofs and empirical results on matrix games demonstrate that GVR ensures optimal consistency under sufficient exploration.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Counterfactual Conservative Q Learning for Offline Multi-agent Reinforcement LearningJianzhun Shao, Yun Qu, Chen Chen, Hongchang Zhang 等NeurIPS 2023 · 被引用 56 次
- Automatic Grouping for Efficient Cooperative Multi-Agent Reinforcement LearningYifan Zang, Jinmin He, Kai Li, Haobo Fu 等NeurIPS 2023 · 被引用 37 次
- Complementary Attention for Multi-Agent Reinforcement LearningJianzhun Shao, Hongchang Zhang, Yun Qu, Chang Liu 等ICML 2023 · 被引用 17 次
- Retaining Suboptimal Actions to Follow Shifting Optima in Multi-Agent Reinforcement LearningYonghyeon Jo, Sunwoo Lee, Seungyul HanICLR 2026 · 被引用 5 次
- Backpropagation Through AgentsZhiyuan Li, Wenshuai Zhao, Lijun Wu, Joni PajarinenAAAI 2024 · 被引用 3 次
它引用的顶会 Paper4
- Weighted QMIX: Expanding Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement LearningTabish Rashid, Gregory Farquhar, Bei Peng, Shimon WhitesonNeurIPS 2020 · 被引用 1,960 次
- Deep Coordination GraphsWendelin Boehmer, Vitaly Kurin, Shimon WhitesonICML 2020 · 被引用 209 次
- Q-value Path Decomposition for Deep Multiagent Reinforcement LearningYaodong Yang, Jianye Hao, Guangyong Chen, Hongyao Tang 等ICML 2020 · 被引用 64 次
- UneVEn: Universal Value Exploration for Multi-Agent Reinforcement LearningTarun Gupta, Anuj Mahajan, Bei Peng, Wendelin Boehmer 等ICML 2021 · 被引用 59 次
相关 Paper
- Beyond Monotonicity: Revisiting Factorization Principles in Multi-Agent Q-LearningTianmeng Hu, Yongzheng Cui, Rui Tang, Biao Luo 等AAAI 2026
- Non-Linear Coordination GraphsYipeng Kang, Tonghan Wang, Qianlan Yang, Xiaoran Wu 等NeurIPS 2022 · 被引用 14 次
- ResQ: A Residual Q Function-based Approach for Multi-Agent Reinforcement Learning Value FactorizationSiqi Shen, Mengwei Qiu, Jun Liu, Weiquan Liu 等NeurIPS 2022 · 被引用 35 次
- Variational Empowerment as Representation Learning for Goal-Conditioned Reinforcement LearningJongwook Choi, Archit Sharma, Honglak Lee, Sergey Levine 等ICML 2021 · 被引用 41 次
- Potentially Optimal Joint Actions Recognition for Cooperative Multi-Agent Reinforcement LearningChang Huang, Shatong Zhu, Junqiao Zhao, Hongtu Zhou 等ICLR 2026
