Shapley Counterfactual Credits for Multi-Agent Reinforcement Learning
Jiahui Li, Kun Kuang, Baoxiang Wang, Furui Liu, Long Chen, Fei Wu, Jun Xiao
Abstract
Centralized Training with Decentralized Execution (CTDE) has been a popular paradigm in cooperative Multi-Agent Reinforcement Learning (MARL) settings and is widely used in many real applications. One of the major challenges in the training process is credit assignment, which aims to deduce the contributions of each agent according to the global rewards. Existing credit assignment methods focus on either decomposing the joint value function into individual value functions or measuring the impact of local observations and actions on the global value function. These approaches lack a thorough consideration of the complicated interactions among multiple agents, leading to an unsuitable assignment of credit and subsequently mediocre results on MARL. We propose Shapley Counterfactual Credit Assignment, a novel method for explicit credit assignment which accounts for the coalition of agents. Specifically, Shapley Value and its desired properties are leveraged in deep MARL to credit any combinations of agents, which grants us the capability to estimate the individual credit for each agent. Despite this capability, the main technical difficulty lies in the computational complexity of Shapley Value who grows factorially as the number of agents. We instead utilize an approximation method via Monte Carlo sampling, which reduces the sample complexity while maintaining its effectiveness. We evaluate our method on StarCraft II benchmarks across different scenarios. Our method outperforms existing cooperative MARL algorithms significantly and achieves the state-of-the-art, with especially large margins on tasks with more severe difficulties.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fc56e6f0-2c8c-4131-b77d-313d026d239cCited by top-tier papers22
- WeightedSHAP: analyzing and improving Shapley based feature attributionsYongchan Kwon, James Y. ZouNeurIPS 2022 · 60 citations
- G2L: Semantically Aligned and Uniform Video Grounding via Geodesic and Game TheoryHongxiang Li, Meng Cao, Xuxin Cheng, Yaowei Li et al.ICCV 2023 · 32 citations
- Deconfounded Value Decomposition for Multi-Agent Reinforcement LearningJiahui Li, Kun Kuang, Baoxiang Wang, Furui Liu et al.ICML 2022 · 28 citations
- NA2Q: Neural Attention Additive Model for Interpretable Multi-Agent Q-LearningZichuan Liu, Yuanyang Zhu, Chunlin ChenICML 2023 · 25 citations
- Instance-wise or Class-wise? A Tale of Neighbor Shapley for Concept-based ExplanationJiahui Li, Kun Kuang, Lin Li, Long Chen et al.ACM MM 2021 · 17 citations
Builds on8
- The Many Shapley Values for Model ExplanationMukund Sundararajan, Amir NajmiICML 2020 · 799 citations
- Problems with Shapley-value-based explanations as feature importance measuresI. Elizabeth Kumar, Suresh Venkatasubramanian, Carlos Scheidegger, Sorelle A. FriedlerICML 2020 · 458 citations
- Asymmetric Shapley values: incorporating causal knowledge into model-agnostic explainabilityChristopher Frye, Colin Rowat, Ilya FeigeNeurIPS 2020 · 246 citations
- Causal Shapley Values: Exploiting Causal Knowledge to Explain Individual Predictions of Complex ModelsTom Heskes, Evi Sijben, Ioan Gabriel Bucur, Tom ClaassenNeurIPS 2020 · 235 citations
- Counterfactual Critic Multi-Agent Training for Scene Graph GenerationLong Chen, Hanwang Zhang, Jun Xiao, Xiangnan He et al.ICCV 2019 · 165 citations
Related papers
- STAS: Spatial-Temporal Return Decomposition for Solving Sparse Rewards Problems in Multi-agent Reinforcement LearningSirui Chen, Zhaowei Zhang, Yaodong Yang, Yali DuAAAI 2024 · 11 citations
- Q-value Path Decomposition for Deep Multiagent Reinforcement LearningYaodong Yang, Jianye Hao, Guangyong Chen, Hongyao Tang et al.ICML 2020 · 64 citations
- Towards Understanding Cooperative Multi-Agent Q-Learning with Value FactorizationJianhao Wang, Zhizhou Ren, Beining Han, Jianing Ye et al.NeurIPS 2021 · 50 citations
- Shapley Q-Value: A Local Reward Approach to Solve Global Reward GamesJianhong Wang, Yuan Zhang, Tae-Kyun Kim, Yunjie GuAAAI 2020 · 159 citations
- Learning Implicit Credit Assignment for Cooperative Multi-Agent Reinforcement LearningMeng Zhou, Ziyu Liu, Pengwei Sui, Yixuan Li et al.NeurIPS 2020 · 142 citations
