SHAQ: Incorporating Shapley Value Theory into Multi-Agent Q-Learning
Jianhong Wang, Yuan Zhang, Yunjie Gu, Tae-Kyun Kim
Abstract
Value factorisation is a useful technique for multi-agent reinforcement learning (MARL) in global reward game, however, its underlying mechanism is not yet fully understood. This paper studies a theoretical framework for value factorisation with interpretability via Shapley value theory. We generalise Shapley value to Markov convex game called Markov Shapley value (MSV) and apply it as a value factorisation method in global reward game, which is obtained by the equivalence between the two games. Based on the properties of MSV, we derive Shapley-Bellman optimality equation (SBOE) to evaluate the optimal MSV, which corresponds to an optimal joint deterministic policy. Furthermore, we propose Shapley-Bellman operator (SBO) that is proved to solve SBOE. With a stochastic approximation and some transformations, a new MARL algorithm called Shapley Q-learning (SHAQ) is established, the implementation of which is guided by the theoretical results of SBO and MSV. We also discuss the relationship between SHAQ and relevant value factorisation methods. In the experiments, SHAQ exhibits not only superior performances on all tasks but also the interpretability that agrees with the theoretical analysis. The implementation of this paper is placed on https://github.com/hsvgbkhgbv/shapley-q-learning .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- One Sample Fits All: Approximating All Probabilistic Values Simultaneously and EfficientlyWeida Li, Yaoliang YuNeurIPS 2024 · 14 citations
- Effective Policy Learning for Multi-Agent Online Coordination Beyond Submodular ObjectivesQixin Zhang, Yan Sun, Can Jin, Xikun Zhang et al.NeurIPS 2025 · 4 citations
- Select to Perfect: Imitating desired behavior from large multi-agent dataTim Franzmeyer, Edith Elkind, Philip Torr, Jakob Nicolaus Foerster et al.ICLR 2024 · 3 citations
- High-order Interactions Modeling for Interpretable Multi-Agent Q-LearningQinyu Xu, Yuanyang Zhu, Xuefei Wu, Chunlin ChenNeurIPS 2025 · 2 citations
- Learning the Expected Core of Strictly Convex Stochastic Cooperative GamesNam Phuong Tran, The-Anh Ta, Shuqing Shi, Debmalya Mandal et al.NeurIPS 2024 · 1 citation
Builds on4
- Weighted QMIX: Expanding Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement LearningTabish Rashid, Gregory Farquhar, Bei Peng, Shimon WhitesonNeurIPS 2020 · 1,960 citations
- Multi-Agent Reinforcement Learning for Active Voltage Control on Power Distribution NetworksJianhong Wang, Wangkun Xu, Yunjie Gu, Wenbin Song et al.NeurIPS 2021 · 216 citations
- Deep Coordination GraphsWendelin Boehmer, Vitaly Kurin, Shimon WhitesonICML 2020 · 209 citations
- Shapley Q-Value: A Local Reward Approach to Solve Global Reward GamesJianhong Wang, Yuan Zhang, Tae-Kyun Kim, Yunjie GuAAAI 2020 · 159 citations
Related papers
- Explaining Reinforcement Learning with Shapley ValuesDaniel Beechey, Thomas M. S. Smith, Özgür SimsekICML 2023 · 41 citations
- ResQ: A Residual Q Function-based Approach for Multi-Agent Reinforcement Learning Value FactorizationSiqi Shen, Mengwei Qiu, Jun Liu, Weiquan Liu et al.NeurIPS 2022 · 35 citations
- DFAC Framework: Factorizing the Value Function via Quantile Mixture for Multi-Agent Distributional Q-LearningWei-Fang Sun, Cheng-Kuang Lee, Chun-Yi LeeICML 2021 · 56 citations
- Towards Understanding Cooperative Multi-Agent Q-Learning with Value FactorizationJianhao Wang, Zhizhou Ren, Beining Han, Jianing Ye et al.NeurIPS 2021 · 50 citations
- FOP: Factorizing Optimal Joint Policy of Maximum-Entropy Multi-Agent Reinforcement LearningTianhao Zhang, Yueheng Li, Chen Wang, Guangming Xie et al.ICML 2021 · 88 citations
