Lune

AAAI2020顶会

SMIX(λ): Enhancing Centralized Value Functions for Cooperative Multi-Agent Reinforcement Learning

Chao Wen, Xinghu Yao, Yuhui Wang, Xiaoyang Tan

2020年份
57被引次数
2顶会引用

摘要

Learning a stable and generalizable centralized value function (CVF) is a crucial but challenging task in multiagent reinforcement learning (MARL), as it has to deal with the issue that the joint action space increases exponentially with the number of agents in such scenarios. This article proposes an approach, named SMIX(<inline-formula> <tex-math notation="LaTeX">λ{\lambda } </tex-math></inline-formula>), that uses an OFF-policy training to achieve this by avoiding the greedy assumption commonly made in CVF learning. As importance sampling for such OFF-policy training is both computationally costly and numerically unstable, we proposed to use the <inline-formula> <tex-math notation="LaTeX">λ{\lambda } </tex-math></inline-formula>-return as a proxy to compute the temporal difference (TD) error. With this new loss function objective, we adopt a modified QMIX network structure as the base to train our model. By further connecting it with the <inline-formula> <tex-math notation="LaTeX">Q(λ){Q(\lambda)} </tex-math></inline-formula> approach from a unified expectation correction viewpoint, we show that the proposed SMIX(<inline-formula> <tex-math notation="LaTeX">λ{\lambda } </tex-math></inline-formula>) is equivalent to <inline-formula> <tex-math notation="LaTeX">Q(λ){Q(\lambda)} </tex-math></inline-formula> and hence shares its convergence properties, while without being suffered from the aforementioned curse of dimensionality problem inherent in MARL. Experiments on the StarCraft Multiagent Challenge (SMAC) benchmark demonstrate that our approach not only outperforms several state-of-the-art MARL methods by a large margin but also can be used as a general tool to improve the overall performance of other centralized training with decentralized execution (CTDE)-type algorithms by enhancing their CVFs.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper2

问问它们各自怎么用它

它引用的顶会 Paper1

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖