SMIX(λ): Enhancing Centralized Value Functions for Cooperative Multi-Agent Reinforcement Learning
Chao Wen, Xinghu Yao, Yuhui Wang, Xiaoyang Tan
摘要
Learning a stable and generalizable centralized value function (CVF) is a crucial but challenging task in multiagent reinforcement learning (MARL), as it has to deal with the issue that the joint action space increases exponentially with the number of agents in such scenarios. This article proposes an approach, named SMIX(<inline-formula> <tex-math notation="LaTeX"> </tex-math></inline-formula>), that uses an OFF-policy training to achieve this by avoiding the greedy assumption commonly made in CVF learning. As importance sampling for such OFF-policy training is both computationally costly and numerically unstable, we proposed to use the <inline-formula> <tex-math notation="LaTeX"> </tex-math></inline-formula>-return as a proxy to compute the temporal difference (TD) error. With this new loss function objective, we adopt a modified QMIX network structure as the base to train our model. By further connecting it with the <inline-formula> <tex-math notation="LaTeX"> </tex-math></inline-formula> approach from a unified expectation correction viewpoint, we show that the proposed SMIX(<inline-formula> <tex-math notation="LaTeX"> </tex-math></inline-formula>) is equivalent to <inline-formula> <tex-math notation="LaTeX"> </tex-math></inline-formula> and hence shares its convergence properties, while without being suffered from the aforementioned curse of dimensionality problem inherent in MARL. Experiments on the StarCraft Multiagent Challenge (SMAC) benchmark demonstrate that our approach not only outperforms several state-of-the-art MARL methods by a large margin but also can be used as a general tool to improve the overall performance of other centralized training with decentralized execution (CTDE)-type algorithms by enhancing their CVFs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Towards a Standardised Performance Evaluation Protocol for Cooperative MARLRihab Gorsane, Omayma Mahjoub, Ruan de Kock, Roland Dubb 等NeurIPS 2022 · 被引用 79 次
- Coordination Between Individual Agents in Multi-Agent Reinforcement LearningYang Zhang, Qingyu Yang, Dou An, Chengwei ZhangAAAI 2021 · 被引用 21 次
它引用的顶会 Paper1
相关 Paper
- Inverse Factorized Soft Q-Learning for Cooperative Multi-agent Imitation LearningThe Viet Bui, Tien Mai, Thanh Hong NguyenNeurIPS 2024 · 被引用 10 次
- Weighted QMIX: Expanding Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement LearningTabish Rashid, Gregory Farquhar, Bei Peng, Shimon WhitesonNeurIPS 2020 · 被引用 1,960 次
- RMIX: Learning Risk-Sensitive Policies for Cooperative Reinforcement Learning AgentsWei Qiu, Xinrun Wang, Runsheng Yu, Rundong Wang 等NeurIPS 2021 · 被引用 71 次
- Shapley Counterfactual Credits for Multi-Agent Reinforcement LearningJiahui Li, Kun Kuang, Baoxiang Wang, Furui Liu 等KDD 2021 · 被引用 49 次
- Value-Decomposition Multi-Agent Actor-CriticsJianyu Su, Stephen C. Adams, Peter A. BelingAAAI 2021 · 被引用 140 次
