Diff-MoE: Diffusion Transformer with Time-Aware and Space-Adaptive Experts
Kun Cheng, Xiao He, Lei Yu, Zhijun Tu, Mingrui Zhu, Nannan Wang, Xinbo Gao, Jie Hu
摘要
Diffusion models have transformed generative modeling but suffer from scalability limitations due to computational overhead and inflexible architectures that process all generative stages and tokens uniformly. In this work, we introduce Diff-MoE, a novel framework that combines Diffusion Transformers with Mixture-of-Experts to exploit both temporarily adaptability and spatial flexibility. Our design incorporates expert-specific timestep conditioning, allowing each expert to process different spatial tokens while adapting to the generative stage, to dynamically allocate resources based on both the temporal and spatial characteristics of the generative task. Additionally, we propose a globally-aware feature recalibration mechanism that amplifies the representational capacity of expert modules by dynamically adjusting feature contributions based on input relevance. Extensive experiments on image generation benchmarks demonstrate that Diff-MoE significantly outperforms state-of-theart methods. Our work demonstrates the potential of integrating diffusion models with expert-based designs, offering a scalable and effective framework for advanced generative modeling. The code is available at https://github.com/ kunncheng/Diff-MoE .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Routing Matters in MoE: Scaling Diffusion Transformers with Explicit Routing GuidanceYujie Wei, Shiwei Zhang, Hangjie Yuan, Yujin Han 等ICLR 2026 · 被引用 26 次
- Elastic Diffusion TransformerJiangshan Wang, Zeqiang Lai, Jiarui Chen, Jiayi Guo 等ICML 2026 · 被引用 7 次
- Mixture of Ranks with Degradation-Aware Routing for One-Step Real-World Image Super-ResolutionXiao He, Zhijun Tu, Kun Cheng, Mingrui Zhu 等AAAI 2026 · 被引用 1 次
- Towards a Unified Generative Model for Scarce Time Series with Domain ExpertsZihao Yao, Qi Zheng, Jiankai Zuo, YAYING ZHANGICML 2026
- Test-Time Instance-Specific Parameter Composition: A New Paradigm for Adaptive Generative ModelingMinh-Tuan Tran, Xuan-May Le, Quan Hung Tran, Mehrtash Harandi 等CVPR 2026
它引用的顶会 Paper23
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
相关 Paper
- Expert Race: A Flexible Routing Strategy for Scaling Diffusion Transformer with Mixture of ExpertsYike Yuan, Ziyu Wang, Zihao Huang, Defa Zhu 等ICML 2025
- EC-DIT: Scaling Diffusion Transformers with Adaptive Expert-Choice RoutingHaotian Sun, Tao Lei, Bowen Zhang, Yanghao Li 等ICLR 2025
- Dense2MoE: Restructuring Diffusion Transformer to MoE for Efficient Text-to-Image GenerationYouwei Zheng, Yuxi Ren, Xin Xia, Xuefeng Xiao 等ICCV 2025 · 被引用 1 次
- DTop-p MoE: Sparsity-Controlled Dynamic Top-p MoE for Foundation Model Pre-trainingCan Jin, Hongwu Peng, Mingcan Xiang, Qixin Zhang 等ICML 2026 · 被引用 3 次
- Diff-MoE: Efficient Batched MoE Inference with Priority-Driven Differential Expert CachingKexin Li, Wenkan Huang, Qinggang Wang, Long Zheng 等SC 2025 · 被引用 3 次
