Dense Backpropagation Improves Training for Sparse Mixture-of-Experts
Ashwinee Panda, Vatsal Baherwani, Zain Sarwar, Benjamin Thérien, Sambit Sahu, Tom Goldstein, Supriyo Chakraborty
Abstract
Mixture of Experts (MoE) pretraining is more scalable than dense Transformer pretraining, because MoEs learn to route inputs to a sparse set of their feedforward parameters. However, this means that MoEs only receive a sparse backward update, leading to training instability and suboptimal performance. We present a lightweight approximation method that gives the MoE router a dense gradient update while continuing to sparsely activate its parameters. Our method, which we refer to as Default MoE, substitutes missing expert activations with default outputs consisting of an exponential moving average of expert outputs previously seen over the course of training. This allows the router to receive signals from every expert for each token, leading to significant improvements in training performance. Our Default MoE outperforms standard TopK routing in a variety of settings without requiring significant computational overhead. Code: https://github.com/vatsal0/default-moe.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- What Layers When: Learning to Skip Compute in LLMs with Residual GatesFilipe Laitenberger, Dawid Jan Kopiczko, Cees G. M. Snoek, Yuki M. AsanoICLR 2026 · 6 citations
- Routing Manifold Alignment Improves Generalization of Mixture-of-Experts LLMsZhongyang Li, Ziyue Li, Tianyi ZhouICLR 2026 · 5 citations
Builds on5
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- GLaM: Efficient Scaling of Language Models with Mixture-of-ExpertsNan Du, Yanping Huang, Andrew M. Dai, Simon Tong et al.ICML 2022 · 1,173 citations
- Unified Scaling Laws for Routed Language ModelsAidan Clark, Diego de Las Casas, Aurelia Guy, Arthur Mensch et al.ICML 2022 · 266 citations
- Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert ModelsZihan Qiu, Zeyu Huang, Bo Zheng, Kaiyue Wen et al.ACL 2025 · 42 citations
- ReMoE: Fully Differentiable Mixture-of-Experts with ReLU RoutingZiteng Wang, Jun Zhu, Jianfei ChenICLR 2025
Related papers
- StableMoE: Stable Routing Strategy for Mixture of ExpertsDamai Dai, Li Dong, Shuming Ma, Bo Zheng et al.ACL 2022
- Teacher-Guided Routing for Sparse Vision Mixture-of-ExpertsMasahiro Kada, Ryota Yoshihashi, Satoshi Ikehata, Rei Kawakami et al.CVPR 2026
- Mixture-of-Experts with Expert Choice RoutingYanqi Zhou, Tao Lei, Hanxiao Liu, Nan Du et al.NeurIPS 2022 · 933 citations
- Mixture of Tokens: Continuous MoE through Cross-Example AggregationSzymon Antoniak, Michal Krutul, Maciej Pióro, Jakub Krajewski et al.NeurIPS 2024 · 6 citations
- DTop-p MoE: Sparsity-Controlled Dynamic Top-p MoE for Foundation Model Pre-trainingCan Jin, Hongwu Peng, Mingcan Xiang, Qixin Zhang et al.ICML 2026 · 3 citations
