Variational Distillation of Diffusion Policies into Mixture of Experts
Hongyi Zhou, Denis Blessing, Ge Li, Onur Celik, Xiaogang Jia, Gerhard Neumann, Rudolf Lioutikov
摘要
This work introduces Variational Diffusion Distillation (VDD), a novel method that distills denoising diffusion policies into Mixtures of Experts (MoE) through variational inference. Diffusion Models are the current state-of-the-art in generative modeling due to their exceptional ability to accurately learn and represent complex, multi-modal distributions. This ability allows Diffusion Models to replicate the inherent diversity in human behavior, making them the preferred models in behavior learning such as Learning from Human Demonstrations (LfD). However, diffusion models come with some drawbacks, including the intractability of likelihoods and long inference times due to their iterative sampling process. The inference times, in particular, pose a significant challenge to real-time applications such as robot control. In contrast, MoEs effectively address the aforementioned issues while retaining the ability to represent complex distributions but are notoriously difficult to train. VDD is the first method that distills pre-trained diffusion models into MoE models, and hence, combines the expressiveness of Diffusion Models with the benefits of Mixture Models. Specifically, VDD leverages a decompositional upper bound of the variational objective that allows the training of each expert separately, resulting in a robust optimization scheme for MoEs. VDD demonstrates across nine complex behavior learning tasks, that it is able to: i) accurately distill complex distributions learned by the diffusion model, ii) outperform existing state-of-the-art distillation methods, and iii) surpass conventional methods for training MoE.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Sparse ActionGen: Accelerating Diffusion Policy with Real-time PruningKangye Ji, Jianbo Zhou, Yuan Meng, Ye Li 等ICML 2026 · 被引用 4 次
- Test-time Sparsity for Extreme Fast Action DiffusionKangye Ji, Yuan Meng, Jianbo Zhou, Ye Li 等CVPR 2026 · 被引用 1 次
- Trust-Region Diffusion Policies for Massively Parallel On-Policy RLHuy Le, Onur Celik, Denis Blessing, Tai Hoang 等ICML 2026
- Beyond Independence: Learning Correlated Views for Variational Incomplete Multi-View ClusteringZheming Xu, Aiyue Tang, Shidi Chen, Xuechao Zou 等ICML 2026
- Aggregation of Dependent Expert Distributions in Multimodal Variational AutoencodersRogelio Andrade Mancisidor, Robert Jenssen, Shujian Yu, Michael KampffmeyerICML 2025
它引用的顶会 Paper28
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion ModelsAlexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam 等ICML 2022 · 被引用 4,691 次
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 被引用 3,959 次
- Consistency ModelsYang Song, Prafulla Dhariwal, Mark Chen, Ilya SutskeverICML 2023 · 被引用 1,720 次
- Improved Techniques for Training Score-Based Generative ModelsYang Song, Stefano ErmonNeurIPS 2020 · 被引用 1,527 次
相关 Paper
- One-Step Diffusion Policy: Fast Visuomotor Policies via Diffusion DistillationZhendong Wang, Max Li, Ajay Mandlekar, Zhenjia Xu 等ICML 2025
- Abstracting Robot Manipulation Skills via Mixture-of-Experts Diffusion PoliciesCe Hao, Xuanran Zhai, Yaohua Liu, Harold SohICLR 2026 · 被引用 14 次
- MotionDiffuser: Controllable Multi-Agent Motion Prediction Using DiffusionChiyu Max Jiang, Andre Cornman, Cheolho Park, Benjamin Sapp 等CVPR 2023
- DiTEA: Mixture-of-Experts for Vision-Language-Action Model in Robotic ManipulationChengxuan Li, Xingwan WangAAAI 2026
- Distillation of Discrete Diffusion through Dimensional CorrelationsSatoshi Hayakawa, Yuhta Takida, Masaaki Imaizumi, Hiromi Wakaki 等ICML 2025
