Variational Distillation of Diffusion Policies into Mixture of Experts
Hongyi Zhou, Denis Blessing, Ge Li, Onur Celik, Xiaogang Jia, Gerhard Neumann, Rudolf Lioutikov
Abstract
This work introduces Variational Diffusion Distillation (VDD), a novel method that distills denoising diffusion policies into Mixtures of Experts (MoE) through variational inference. Diffusion Models are the current state-of-the-art in generative modeling due to their exceptional ability to accurately learn and represent complex, multi-modal distributions. This ability allows Diffusion Models to replicate the inherent diversity in human behavior, making them the preferred models in behavior learning such as Learning from Human Demonstrations (LfD). However, diffusion models come with some drawbacks, including the intractability of likelihoods and long inference times due to their iterative sampling process. The inference times, in particular, pose a significant challenge to real-time applications such as robot control. In contrast, MoEs effectively address the aforementioned issues while retaining the ability to represent complex distributions but are notoriously difficult to train. VDD is the first method that distills pre-trained diffusion models into MoE models, and hence, combines the expressiveness of Diffusion Models with the benefits of Mixture Models. Specifically, VDD leverages a decompositional upper bound of the variational objective that allows the training of each expert separately, resulting in a robust optimization scheme for MoEs. VDD demonstrates across nine complex behavior learning tasks, that it is able to: i) accurately distill complex distributions learned by the diffusion model, ii) outperform existing state-of-the-art distillation methods, and iii) surpass conventional methods for training MoE.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 57aeb724-e57b-4781-b4c9-b330a6f6584fCited by top-tier papers6
- Sparse ActionGen: Accelerating Diffusion Policy with Real-time PruningKangye Ji, Jianbo Zhou, Yuan Meng, Ye Li et al.ICML 2026 · 4 citations
- Test-time Sparsity for Extreme Fast Action DiffusionKangye Ji, Yuan Meng, Jianbo Zhou, Ye Li et al.CVPR 2026 · 1 citation
- Trust-Region Diffusion Policies for Massively Parallel On-Policy RLHuy Le, Onur Celik, Denis Blessing, Tai Hoang et al.ICML 2026
- Beyond Independence: Learning Correlated Views for Variational Incomplete Multi-View ClusteringZheming Xu, Aiyue Tang, Shidi Chen, Xuechao Zou et al.ICML 2026
- Aggregation of Dependent Expert Distributions in Multimodal Variational AutoencodersRogelio Andrade Mancisidor, Robert Jenssen, Shujian Yu, Michael KampffmeyerICML 2025
Builds on28
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion ModelsAlexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam et al.ICML 2022 · 4,691 citations
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 3,959 citations
- Consistency ModelsYang Song, Prafulla Dhariwal, Mark Chen, Ilya SutskeverICML 2023 · 1,720 citations
- Improved Techniques for Training Score-Based Generative ModelsYang Song, Stefano ErmonNeurIPS 2020 · 1,527 citations
Related papers
- One-Step Diffusion Policy: Fast Visuomotor Policies via Diffusion DistillationZhendong Wang, Max Li, Ajay Mandlekar, Zhenjia Xu et al.ICML 2025
- Abstracting Robot Manipulation Skills via Mixture-of-Experts Diffusion PoliciesCe Hao, Xuanran Zhai, Yaohua Liu, Harold SohICLR 2026 · 14 citations
- MotionDiffuser: Controllable Multi-Agent Motion Prediction Using DiffusionChiyu Max Jiang, Andre Cornman, Cheolho Park, Benjamin Sapp et al.CVPR 2023
- DiTEA: Mixture-of-Experts for Vision-Language-Action Model in Robotic ManipulationChengxuan Li, Xingwan WangAAAI 2026
- Distillation of Discrete Diffusion through Dimensional CorrelationsSatoshi Hayakawa, Yuhta Takida, Masaaki Imaizumi, Hiromi Wakaki et al.ICML 2025
