Discovering Generalizable Multi-agent Coordination Skills from Multi-task Offline Data
Fuxiang Zhang, Chengxing Jia, Yi-Chen Li, Lei Yuan, Yang Yu, Zongzhang Zhang
Abstract
Cooperative multi-agent reinforcement learning (MARL) faces the challenge of adapting to multiple tasks with varying agents and targets. Previous multi-task MARL approaches require costly interactions to simultaneously learn or fine-tune policies in different tasks. However, the situation that an agent should generalize to multiple tasks with only offline data from limited tasks is more in line with the needs of real-world applications. Since offline multi-task data contains a variety of behaviors, an effective data-driven approach is to extract informative latent variables that can represent universal skills for realizing coordination across tasks. In this paper, we propose a novel Offline MARL algorithm to Discover coordInation Skills (ODIS) from multi-task data. ODIS first extracts task-invariant coordination skills from offline multi-task data and learns to delineate different agent behaviors with the discovered coordination skills. Then we train a coordination policy to choose optimal coordination skills with the centralized training and decentralized execution paradigm. We further demonstrate that the discovered coordination skills can assign effective coordinative behaviors, thus significantly enhancing generalization to unseen tasks. Empirical results in cooperative MARL benchmarks, including the StarCraft multi-agent challenge, show that ODIS obtains superior performance in a wide range of tasks only with offline data from limited sources.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers16
- Hierarchical Multi-Agent Skill DiscoveryMingyu Yang, Yaodong Yang, Zhenbo Lu, Wengang Zhou et al.NeurIPS 2023 · 34 citations
- Decompose a Task into Generalizable Subtasks in Multi-Agent Reinforcement LearningZikang Tian, Ruizhi Chen, Xing Hu, Ling Li et al.NeurIPS 2023 · 23 citations
- Learning and Planning Multi-Agent Tasks via an MoE-based World ModelZijie Zhao, Zhongyue Zhao, Kaixuan Xu, Yuqian Fu et al.NeurIPS 2025 · 12 citations
- A Unified Algorithm Framework for Unsupervised Discovery of Skills based on Determinantal Point ProcessJiayu Chen, Vaneet Aggarwal, Tian LanNeurIPS 2023 · 8 citations
- Multi-Agent Guided Policy OptimizationYueheng Li, Guangming Xie, Zongqing LuICLR 2026 · 4 citations
Related papers
- Learning Generalizable Skills from Offline Multi-Task Data for Multi-Agent CooperationSicong Liu, Yang Shu, Chenjuan Guo, Bin YangICLR 2025
- Goal-Oriented Skill Abstraction for Offline Multi-Task Reinforcement LearningJinmin He, Kai Li, Yifan Zang, Haobo Fu et al.ICML 2025
- Conservative Data Sharing for Multi-Task Offline Reinforcement LearningTianhe Yu, Aviral Kumar, Yevgen Chebotar, Karol Hausman et al.NeurIPS 2021 · 94 citations
- MangoBench: A Benchmark for Multi-Agent Goal-Conditioned Offline Reinforcement LearningYi Wang, Ningze Zhong, Zhiheng Fu, Longguang Wang et al.CVPR 2026
- Offline Multi-Agent Reinforcement Learning via In-Sample Sequential Policy OptimizationZongkai Liu, Qian Lin, Chao Yu, Xiawei Wu et al.AAAI 2025 · 3 citations
