Expert Race: A Flexible Routing Strategy for Scaling Diffusion Transformer with Mixture of Experts
Yike Yuan, Ziyu Wang, Zihao Huang, Defa Zhu, Xun Zhou, Jingyi Yu, Qiyang Min
Abstract
Diffusion models have emerged as mainstream framework in visual generation. Building upon this success, the integration of Mixture of Experts (MoE) methods has shown promise in enhancing model scalability and performance. In this paper, we introduce Race-DiT, a novel MoE model for diffusion transformers with a flexible routing strategy, Expert Race. By allowing tokens and experts to compete together and select the top candidates, the model learns to dynamically assign experts to critical tokens. Additionally, we propose per-layer regularization to address challenges in shallow layer learning, and router similarity loss to prevent mode collapse, ensuring better expert utilization. Extensive experiments on ImageNet validate the effectiveness of our approach, showcasing significant performance gains while promising scaling properties.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e6944ad9-db58-4448-abbc-39bfae6a11bdCited by top-tier papers4
- Routing Matters in MoE: Scaling Diffusion Transformers with Explicit Routing GuidanceYujie Wei, Shiwei Zhang, Hangjie Yuan, Yujin Han et al.ICLR 2026 · 26 citations
- UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context LearningZihao Huang, Yu Bao, Qiyang Min, Siyan Chen et al.ICLR 2026 · 6 citations
- Rethinking Sparse Mixture of Experts from a Unified PerspectiveGiang Do, Hung Le, Truyen TranICML 2026 · 1 citation
- MoEG-HOI: Mixture of Expert Groups for One-Stage Hand-Object Interaction Motion Generation with Hand-Finger-Joint Semantic GuidanceHang Xu, Yang Xiao, Changlong Jiang, Haohong Kuang et al.AAAI 2026
Builds on27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- Diff-MoE: Diffusion Transformer with Time-Aware and Space-Adaptive ExpertsKun Cheng, Xiao He, Lei Yu, Zhijun Tu et al.ICML 2025
- EC-DIT: Scaling Diffusion Transformers with Adaptive Expert-Choice RoutingHaotian Sun, Tao Lei, Bowen Zhang, Yanghao Li et al.ICLR 2025
- Remix-DiT: Mixing Diffusion Transformers for Multi-Expert DenoisingGongfan Fang, Xinyin Ma, Xinchao WangNeurIPS 2024 · 7 citations
- DTop-p MoE: Sparsity-Controlled Dynamic Top-p MoE for Foundation Model Pre-trainingCan Jin, Hongwu Peng, Mingcan Xiang, Qixin Zhang et al.ICML 2026 · 3 citations
- TAG-MoE: Task-Aware Gating for Unified Generative Mixture-of-ExpertsYu Xu, Hongbin Yan, Juan Cao, Yiji Cheng et al.CVPR 2026 · 7 citations
