SMIX(λ): Enhancing Centralized Value Functions for Cooperative Multi-Agent Reinforcement Learning
Chao Wen, Xinghu Yao, Yuhui Wang, Xiaoyang Tan
Abstract
Learning a stable and generalizable centralized value function (CVF) is a crucial but challenging task in multiagent reinforcement learning (MARL), as it has to deal with the issue that the joint action space increases exponentially with the number of agents in such scenarios. This article proposes an approach, named SMIX(<inline-formula> <tex-math notation="LaTeX"> </tex-math></inline-formula>), that uses an OFF-policy training to achieve this by avoiding the greedy assumption commonly made in CVF learning. As importance sampling for such OFF-policy training is both computationally costly and numerically unstable, we proposed to use the <inline-formula> <tex-math notation="LaTeX"> </tex-math></inline-formula>-return as a proxy to compute the temporal difference (TD) error. With this new loss function objective, we adopt a modified QMIX network structure as the base to train our model. By further connecting it with the <inline-formula> <tex-math notation="LaTeX"> </tex-math></inline-formula> approach from a unified expectation correction viewpoint, we show that the proposed SMIX(<inline-formula> <tex-math notation="LaTeX"> </tex-math></inline-formula>) is equivalent to <inline-formula> <tex-math notation="LaTeX"> </tex-math></inline-formula> and hence shares its convergence properties, while without being suffered from the aforementioned curse of dimensionality problem inherent in MARL. Experiments on the StarCraft Multiagent Challenge (SMAC) benchmark demonstrate that our approach not only outperforms several state-of-the-art MARL methods by a large margin but also can be used as a general tool to improve the overall performance of other centralized training with decentralized execution (CTDE)-type algorithms by enhancing their CVFs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3b09d61a-540c-4e6e-a883-5ae173feb524Cited by top-tier papers2
- Towards a Standardised Performance Evaluation Protocol for Cooperative MARLRihab Gorsane, Omayma Mahjoub, Ruan de Kock, Roland Dubb et al.NeurIPS 2022 · 79 citations
- Coordination Between Individual Agents in Multi-Agent Reinforcement LearningYang Zhang, Qingyu Yang, Dou An, Chengwei ZhangAAAI 2021 · 21 citations
Builds on1
Related papers
- Inverse Factorized Soft Q-Learning for Cooperative Multi-agent Imitation LearningThe Viet Bui, Tien Mai, Thanh Hong NguyenNeurIPS 2024 · 10 citations
- Weighted QMIX: Expanding Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement LearningTabish Rashid, Gregory Farquhar, Bei Peng, Shimon WhitesonNeurIPS 2020 · 1,960 citations
- RMIX: Learning Risk-Sensitive Policies for Cooperative Reinforcement Learning AgentsWei Qiu, Xinrun Wang, Runsheng Yu, Rundong Wang et al.NeurIPS 2021 · 71 citations
- Shapley Counterfactual Credits for Multi-Agent Reinforcement LearningJiahui Li, Kun Kuang, Baoxiang Wang, Furui Liu et al.KDD 2021 · 49 citations
- Value-Decomposition Multi-Agent Actor-CriticsJianyu Su, Stephen C. Adams, Peter A. BelingAAAI 2021 · 140 citations
