Conditional Diffusion Model for Multi-Agent Dynamic Task Decomposition
Yanda Zhu, Yuanyang Zhu, Daoyi Dong, Caihua Chen, Chunlin Chen
Abstract
Task decomposition has shown promise in complex cooperative multi-agent reinforcement learning (MARL) tasks, which enables efficient hierarchical learning for long-horizon tasks in dynamic and uncertain environments. However, learning dynamic task decomposition from scratch generally requires a large number of training samples, especially exploring the large joint action space under partial observability. In this paper, we present the Conditional Diffusion Model for Dynamic Task Decomposition (CD 3 T), a novel two-level hierarchical MARL framework designed to automatically infer subtask and coordination patterns. The high-level policy learns subtask representation to generate a subtask selection strategy based on subtask effects. To capture the effects of subtasks on the environment, CD 3 T predicts the next observation and reward using a conditional diffusion model. At the low level, agents collaboratively learn and share specialized skills within their assigned subtasks. Moreover, the learned subtask representation is also used as additional semantic information in a multi-head attention mixing network to enhance value decomposition and provide an efficient reasoning bridge between individual and joint value functions. Experimental results on various benchmarks demonstrate that CD 3 T achieves better performance than existing baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on27
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Weighted QMIX: Expanding Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement LearningTabish Rashid, Gregory Farquhar, Bei Peng, Shimon WhitesonNeurIPS 2020 · 1,960 citations
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar et al.ICLR 2021 · 1,270 citations
- Planning with Diffusion for Flexible Behavior SynthesisMichael Janner, Yilun Du, Joshua B. Tenenbaum, Sergey LevineICML 2022 · 1,115 citations
- Dynamics-Aware Unsupervised Discovery of SkillsArchit Sharma, Shixiang Gu, Sergey Levine, Vikash Kumar et al.ICLR 2020 · 475 citations
Related papers
- Hierarchical Entity-centric Reinforcement Learning with Factored Subgoal DiffusionDan Haramati, Carl Qi, Tal Daniel, Amy Zhang et al.ICLR 2026 · 7 citations
- T4NMTD: Transition-Centric Reinforcement Learning for Non-Markovian Task DecompositionRuixuan Miao, Xu Lu, Cong Tian, Bin Yu et al.AAAI 2026
- Hierarchical Reinforcement Learning with Uncertainty-Guided Diffusional SubgoalsVivienne Huiling Wang, Tinghuai Wang, Joni PajarinenICML 2025
- Hierarchical Multi-Agent Skill DiscoveryMingyu Yang, Yaodong Yang, Zhenbo Lu, Wengang Zhou et al.NeurIPS 2023 · 34 citations
- HAVEN: Hierarchical Cooperative Multi-Agent Reinforcement Learning with Dual Coordination MechanismZhiwei Xu, Yunpeng Bai, Bin Zhang, Dapeng Li et al.AAAI 2023 · 46 citations
