UPDeT: Universal Multi-agent RL via Policy Decoupling with Transformers
Siyi Hu, Fengda Zhu, Xiaojun Chang, Xiaodan Liang
Abstract
Recent advances in multi-agent reinforcement learning have been largely limited training one model from scratch for every new task. This limitation occurs due to the restriction of the model architecture related to fixed input and output dimensions, which hinder the experience accumulation and transfer of the learned agent over tasks across diverse levels of difficulty (e.g. 3 vs 3 or 5 vs 6 multiagent games). In this paper, we make the first attempt to explore a universal multi-agent reinforcement learning pipeline, designing a single architecture to fit tasks with different observation and action configuration requirements. Unlike previous RNN-based models, we utilize a transformer-based model to generate a flexible policy by decoupling the policy distribution from the intertwined input observation, using an importance weight determined with the aid of the selfattention mechanism. Compared to a standard transformer block, the proposed model, which we name Universal Policy Decoupling Transformer (UPDeT), further relaxes the action restriction and makes the multi-agent task's decision process more explainable. UPDeT is general enough to be plugged into any multiagent reinforcement learning pipeline and equip it with strong generalization abilities that enable multiple tasks to be handled at a time. Extensive experiments on large-scale SMAC multi-agent competitive games demonstrate that the proposed UPDeT-based multi-agent reinforcement learning achieves significant improvements relative to state-of-the-art approaches, demonstrating advantageous transfer capability in terms of both performance and training speed (10 times faster). Code is available at https://github.com/hhhusiyi-monash/UPDeT
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers17
- Cooperative Exploration for Multi-Agent Deep Reinforcement LearningIou-Jen Liu, Unnat Jain, Raymond A. Yeh, Alexander G. SchwingICML 2021 · 133 citations
- Randomized Entity-wise Factorization for Multi-Agent Reinforcement LearningShariq Iqbal, Christian A. Schröder de Witt, Bei Peng, Wendelin Boehmer et al.ICML 2021 · 84 citations
- Offline Multi-Agent Reinforcement Learning with Knowledge DistillationWei-Cheng Tseng, Tsun-Hsuan Johnson Wang, Yen-Chen Lin, Phillip IsolaNeurIPS 2022 · 62 citations
- Revisiting Some Common Practices in Cooperative Multi-Agent Reinforcement LearningWei Fu, Chao Yu, Zelai Xu, Jiaqi Yang et al.ICML 2022 · 49 citations
- Self-Organized Group for Cooperative Multi-agent Reinforcement LearningJianzhun Shao, Zhiqiang Lou, Hongchang Zhang, Yuhang Jiang et al.NeurIPS 2022 · 41 citations
Builds on2
Related papers
- Multi-Agent Reinforcement Learning is a Sequence Modeling ProblemMuning Wen, Jakub Grudzien Kuba, Runji Lin, Weinan Zhang et al.NeurIPS 2022 · 408 citations
- Multi-Game Decision TransformersKuang-Huei Lee, Ofir Nachum, Mengjiao Yang, Lisa Lee et al.NeurIPS 2022 · 279 citations
- UniGraspTransformer: Simplified Policy Distillation for Scalable Dexterous Robotic GraspingWenbo Wang, Fangyun Wei, Lei Zhou, Xi Chen et al.CVPR 2025
- STAIRS-Former: Spatio-Temporal Attention with Interleaved Recursive Structure TransFormer for Offline Mulit-task Multi-agent Reinforcement LearningJiwon Jeon, Myungsik Cho, Youngchul SungICLR 2026 · 1 citation
- Transformer-based Working Memory for Multiagent Reinforcement Learning with Action ParsingYaodong Yang, Guangyong Chen, Weixun Wang, Xiaotian Hao et al.NeurIPS 2022 · 24 citations
