UPDeT: Universal Multi-agent RL via Policy Decoupling with Transformers
Siyi Hu, Fengda Zhu, Xiaojun Chang, Xiaodan Liang
摘要
Recent advances in multi-agent reinforcement learning have been largely limited training one model from scratch for every new task. This limitation occurs due to the restriction of the model architecture related to fixed input and output dimensions, which hinder the experience accumulation and transfer of the learned agent over tasks across diverse levels of difficulty (e.g. 3 vs 3 or 5 vs 6 multiagent games). In this paper, we make the first attempt to explore a universal multi-agent reinforcement learning pipeline, designing a single architecture to fit tasks with different observation and action configuration requirements. Unlike previous RNN-based models, we utilize a transformer-based model to generate a flexible policy by decoupling the policy distribution from the intertwined input observation, using an importance weight determined with the aid of the selfattention mechanism. Compared to a standard transformer block, the proposed model, which we name Universal Policy Decoupling Transformer (UPDeT), further relaxes the action restriction and makes the multi-agent task's decision process more explainable. UPDeT is general enough to be plugged into any multiagent reinforcement learning pipeline and equip it with strong generalization abilities that enable multiple tasks to be handled at a time. Extensive experiments on large-scale SMAC multi-agent competitive games demonstrate that the proposed UPDeT-based multi-agent reinforcement learning achieves significant improvements relative to state-of-the-art approaches, demonstrating advantageous transfer capability in terms of both performance and training speed (10 times faster). Code is available at https://github.com/hhhusiyi-monash/UPDeT
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Cooperative Exploration for Multi-Agent Deep Reinforcement LearningIou-Jen Liu, Unnat Jain, Raymond A. Yeh, Alexander G. SchwingICML 2021 · 被引用 133 次
- Randomized Entity-wise Factorization for Multi-Agent Reinforcement LearningShariq Iqbal, Christian A. Schröder de Witt, Bei Peng, Wendelin Boehmer 等ICML 2021 · 被引用 84 次
- Offline Multi-Agent Reinforcement Learning with Knowledge DistillationWei-Cheng Tseng, Tsun-Hsuan Johnson Wang, Yen-Chen Lin, Phillip IsolaNeurIPS 2022 · 被引用 62 次
- Revisiting Some Common Practices in Cooperative Multi-Agent Reinforcement LearningWei Fu, Chao Yu, Zelai Xu, Jiaqi Yang 等ICML 2022 · 被引用 49 次
- Self-Organized Group for Cooperative Multi-agent Reinforcement LearningJianzhun Shao, Zhiqiang Lou, Hongchang Zhang, Yuhang Jiang 等NeurIPS 2022 · 被引用 41 次
它引用的顶会 Paper2
相关 Paper
- Multi-Agent Reinforcement Learning is a Sequence Modeling ProblemMuning Wen, Jakub Grudzien Kuba, Runji Lin, Weinan Zhang 等NeurIPS 2022 · 被引用 408 次
- Multi-Game Decision TransformersKuang-Huei Lee, Ofir Nachum, Mengjiao Yang, Lisa Lee 等NeurIPS 2022 · 被引用 279 次
- UniGraspTransformer: Simplified Policy Distillation for Scalable Dexterous Robotic GraspingWenbo Wang, Fangyun Wei, Lei Zhou, Xi Chen 等CVPR 2025
- STAIRS-Former: Spatio-Temporal Attention with Interleaved Recursive Structure TransFormer for Offline Mulit-task Multi-agent Reinforcement LearningJiwon Jeon, Myungsik Cho, Youngchul SungICLR 2026 · 被引用 1 次
- Transformer-based Working Memory for Multiagent Reinforcement Learning with Action ParsingYaodong Yang, Guangyong Chen, Weixun Wang, Xiaotian Hao 等NeurIPS 2022 · 被引用 24 次
