Transformer-based Working Memory for Multiagent Reinforcement Learning with Action Parsing
Yaodong Yang, Guangyong Chen, Weixun Wang, Xiaotian Hao, Jianye Hao, Pheng-Ann Heng
摘要
Learning in real-world multiagent tasks is challenging due to the usual partial observability of each agent. Previous efforts alleviate the partial observability by historical hidden states with Recurrent Neural Networks, however, they do not consider the multiagent characters that either the multiagent observation consists of a number of object entities or the action space shows clear entity interactions. To tackle these issues, we propose the Agent Transformer Memory (ATM) network with a transformer-based memory. First, ATM utilizes the transformer to enable the unified processing of the factored environmental entities and memory. Inspired by the human’s working memory process where a limited capacity of information temporarily held in mind can effectively guide the decision-making, ATM updates its fixed-capacity memory with the working memory updating schema. Second, as agents’ each action has its particular interaction entities in the environment, ATM parses the action space to introduce this action’s semantic inductive bias by binding each action with its specified involving entity to predict the state-action value or logit. Extensive experiments on the challenging SMAC and Level-Based Foraging environments validate that ATM could boost existing multiagent RL algorithms with impressive learning acceleration and performance improvement.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- RiskQ: Risk-sensitive Multi-Agent Reinforcement Learning Value FactorizationSiqi Shen, Chennan Ma, Chao Li, Weiquan Liu 等NeurIPS 2023 · 被引用 34 次
- Sample-Efficient Multiagent Reinforcement Learning with Reset ReplayYaodong Yang, Guangyong Chen, Jianye Hao, Pheng-Ann HengICML 2024 · 被引用 9 次
- Conditional Diffusion Model for Multi-Agent Dynamic Task DecompositionYanda Zhu, Yuanyang Zhu, Daoyi Dong, Caihua Chen 等AAAI 2026
- Learning Generalizable Skills from Offline Multi-Task Data for Multi-Agent CooperationSicong Liu, Yang Shu, Chenjuan Guo, Bin YangICLR 2025
- Enhancing Cooperative Multi-Agent Reinforcement Learning with State Modelling and Adversarial ExplorationAndreas Kontogiannis, Konstantinos Papathanasiou, Yi Shen, Giorgos Stamou 等ICML 2025
它引用的顶会 Paper11
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee 等NeurIPS 2021 · 被引用 2,557 次
- Offline Reinforcement Learning as One Big Sequence Modeling ProblemMichael Janner, Qiyang Li, Sergey LevineNeurIPS 2021 · 被引用 950 次
- QPLEX: Duplex Dueling Multi-Agent Q-LearningJianhao Wang, Zhizhou Ren, Terry Liu, Yang Yu 等ICLR 2021 · 被引用 595 次
- Stabilizing Transformers for Reinforcement LearningEmilio Parisotto, H. Francis Song, Jack W. Rae, Razvan Pascanu 等ICML 2020 · 被引用 464 次
- Improving Policies via Search in Cooperative Partially Observable GamesAdam Lerer, Hengyuan Hu, Jakob N. Foerster, Noam BrownAAAI 2020 · 被引用 87 次
相关 Paper
- Working Memory GraphsRicky Loynd, Roland Fernandez, Asli Celikyilmaz, Adith Swaminathan 等ICML 2020 · 被引用 41 次
- Memo: Training Memory-Efficient Embodied Agents with Reinforcement LearningGunshi Gupta, Karmesh Yadav, Zsolt Kira, Yarin Gal 等NeurIPS 2025 · 被引用 9 次
- UPDeT: Universal Multi-agent RL via Policy Decoupling with TransformersSiyi Hu, Fengda Zhu, Xiaojun Chang, Xiaodan LiangICLR 2021 · 被引用 49 次
- Learning Multi-Agent Communication through Structured Attentive ReasoningMurtaza Rangwala, Ryan WilliamsNeurIPS 2020 · 被引用 42 次
- Multi-agent In-context Coordination via Decentralized Memory RetrievalTao Jiang, Zichuan Lin, Lihe Li, Yi-Chen Li 等AAAI 2026 · 被引用 1 次
