SkillDiffuser: Interpretable Hierarchical Planning via Skill Abstractions in Diffusion-Based Task Execution
Zhixuan Liang, Yao Mu, Hengbo Ma, Masayoshi Tomizuka, Mingyu Ding, Ping Luo
Abstract
Diffusion models have demonstrated strong potential for robotic trajectory planning. However, generating coherent trajectories from high-level instructions remains challenging, especially for long-range composition tasks requiring multiple sequential skills. We propose SkillDiffuser, an end-to-end hierarchical planning framework integrating interpretable skill learning with conditional diffusion planning to address this problem. At the higher level, the skill abstraction module learns discrete, human-understandable skill representations from visual observations and language instructions. These learned skill embeddings are then used to condition the diffusion model to generate customized latent trajectories aligned with the skills. This allows generating diverse state trajectories that adhere to the learnable skills. By integrating skill learning with conditional trajectory generation, SkillDiffuser produces coherent behavior following abstract instructions across diverse tasks. Experiments on multi-task robotic manipulation benchmarks like Meta-World and LOReL demonstrate state-of-the-art performance and human-interpretable skill representations from SkillDiffuser. More visualization results and information could be found on our website.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 769ad782-0353-41a4-82cb-79eb292b2978Cited by top-tier papers32
- RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic ManipulationTianxing Chen, Zanxin Chen, Baijun Chen, Zijian Cai et al.ICML 2026 · 394 citations
- QueST: Self-Supervised Skill Abstractions for Learning Continuous ControlAtharva Mete, Haotian Xue, Albert Wilcox, Yongxin Chen et al.NeurIPS 2024 · 76 citations
- Learning an Actionable Discrete Diffusion Policy via Large-Scale Actionless Video Pre-TrainingHaoran He, Chenjia Bai, Ling Pan, Weinan Zhang et al.NeurIPS 2024 · 38 citations
- VLMimic: Vision Language Models are Visual Imitation Learner for Fine-grained ActionsGuangyan Chen, Meiling Wang, Te Cui, Yao Mu et al.NeurIPS 2024 · 24 citations
- Diffusion Policy Attacker: Crafting Adversarial Attacks for Diffusion-based PoliciesYipu Chen, Haotian Xue, Yongxin ChenNeurIPS 2024 · 23 citations
Builds on15
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- Planning with Diffusion for Flexible Behavior SynthesisMichael Janner, Yilun Du, Joshua B. Tenenbaum, Sergey LevineICML 2022 · 1,115 citations
- Learning Universal Policies via Text-Guided Video GenerationYilun Du, Sherry Yang, Bo Dai, Hanjun Dai et al.NeurIPS 2023 · 742 citations
- Dynamics-Aware Unsupervised Discovery of SkillsArchit Sharma, Shixiang Gu, Sergey Levine, Vikash Kumar et al.ICLR 2020 · 475 citations
Related papers
- Generative Trajectory Stitching through Diffusion CompositionYunhao Luo, Utkarsh A. Mishra, Yilun Du, Danfei XuNeurIPS 2025 · 48 citations
- Discrete Latent Plans via Semantic Skill AbstractionsHaobin Jiang, Jiangxing Wang, Zongqing LuICLR 2025
- Learning Diffusion Policy from Primitive Skills for Robot ManipulationZhihao Gu, Ming Yang, Difan Zou, Dong XuAAAI 2026
- Pixel Motion Diffusion is What We Need for Robot ControlE-Ro Nguyen, Yichi Zhang, Kanchana Ranasinghe, Xiang Li et al.CVPR 2026 · 10 citations
- Language Control Diffusion: Efficiently Scaling through Space, Time, and TasksEdwin Zhang, Yujie Lu, Shinda Huang, William Yang Wang et al.ICLR 2024 · 34 citations
