Q-value Path Decomposition for Deep Multiagent Reinforcement Learning
Yaodong Yang, Jianye Hao, Guangyong Chen, Hongyao Tang, Yingfeng Chen, Yujing Hu, Changjie Fan, Zhongyu Wei
Abstract
Recently, deep multiagent reinforcement learning (MARL) has become a highly active research area as many real-world problems can be inherently viewed as multiagent systems. A particularly interesting and widely applicable class of problems is the partially observable cooperative multiagent setting, in which a team of agents learns to coordinate their behaviors conditioning on their private observations and commonly shared global reward signals. One natural solution is to resort to the centralized training and decentralized execution paradigm. During centralized training, one key challenge is the multiagent credit assignment: how to allocate the global rewards for individual agent policies for better coordination towards maximizing system-level's benefits. In this paper, we propose a new method called Q-value Path Decomposition (QPD) to decompose the system's global Q-values into individual agents' Q-values. Unlike previous works which restrict the representation relation of the individual Q-values and the global one, we leverage the integrated gradient attribution technique into deep MARL to directly decompose global Q-values along trajectory paths to assign credits for agents. We evaluate QPD on the challenging StarCraft II micromanagement tasks and show that QPD achieves the state-of-the-art performance in both homogeneous and heterogeneous multiagent scenarios compared with existing cooperative MARL algorithms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d28a3583-c798-4861-a229-68466e6dee88Cited by top-tier papers13
- Shapley Counterfactual Credits for Multi-Agent Reinforcement LearningJiahui Li, Kun Kuang, Baoxiang Wang, Furui Liu et al.KDD 2021 · 49 citations
- Automatic Grouping for Efficient Cooperative Multi-Agent Reinforcement LearningYifan Zang, Jinmin He, Kai Li, Haobo Fu et al.NeurIPS 2023 · 37 citations
- Heterogeneous Skill Learning for Multi-agent TasksYuntao Liu, Yuan Li, Xinhai Xu, Yong Dou et al.NeurIPS 2022 · 33 citations
- RACE: Improve Multi-Agent Reinforcement Learning with Representation Asymmetry and Collaborative EvolutionPengyi Li, Jianye Hao, Hongyao Tang, Yan Zheng et al.ICML 2023 · 31 citations
- Deconfounded Value Decomposition for Multi-Agent Reinforcement LearningJiahui Li, Kun Kuang, Baoxiang Wang, Furui Liu et al.ICML 2022 · 28 citations
Related papers
- Learning Explicit Credit Assignment for Cooperative Multi-Agent Reinforcement Learning via Polarization Policy GradientWubing Chen, Wenbin Li, Xiao Liu, Shangdong Yang et al.AAAI 2023 · 11 citations
- DOP: Off-Policy Multi-Agent Decomposed Policy GradientsYihan Wang, Beining Han, Tonghan Wang, Heng Dong et al.ICLR 2021 · 208 citations
- Multiagent Q-learning with Sub-Team CoordinationWenhan Huang, Kai Li, Kun Shao, Tianze Zhou et al.NeurIPS 2022 · 12 citations
- Shapley Q-Value: A Local Reward Approach to Solve Global Reward GamesJianhong Wang, Yuan Zhang, Tae-Kyun Kim, Yunjie GuAAAI 2020 · 159 citations
- STAS: Spatial-Temporal Return Decomposition for Solving Sparse Rewards Problems in Multi-agent Reinforcement LearningSirui Chen, Zhaowei Zhang, Yaodong Yang, Yali DuAAAI 2024 · 11 citations
