ReMoE: Fully Differentiable Mixture-of-Experts with ReLU Routing
Ziteng Wang, Jun Zhu, Jianfei Chen
摘要
Sparsely activated Mixture-of-Experts (MoE) models are widely adopted to scale up model capacity without increasing the computation budget. However, vanilla TopK routers are trained in a discontinuous, non-differentiable way, limiting their performance and scalability. To address this issue, we propose ReMoE, a fully differentiable MoE architecture that offers a simple yet effective drop-in replacement for the conventional TopK+Softmax routing, utilizing ReLU as the router instead. We further propose methods to regulate the router's sparsity while balancing the load among experts. ReMoE's continuous nature enables efficient dynamic allocation of computation across tokens and layers, while also exhibiting domain specialization. Our experiments demonstrate that ReMoE consistently outperforms vanilla TopK-routed MoE across various model sizes, expert counts, and levels of granularity. Furthermore, ReMoE exhibits superior scalability with respect to the number of experts, surpassing traditional MoE architectures. The implementation based on Megatron-LM is available at https://github.com/thu-ml/ReMoE .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- LD-MoLE: Learnable Dynamic Routing for Mixture of LoRA ExpertsYuan Zhuang, Yi Shen, Yuexin Bian, Qing Su 等ICLR 2026 · 被引用 15 次
- Dense Backpropagation Improves Training for Sparse Mixture-of-ExpertsAshwinee Panda, Vatsal Baherwani, Zain Sarwar, Benjamin Thérien 等NeurIPS 2025 · 被引用 10 次
- Analytical FFN-to-MoE Restructuring via Activation Pattern AnalysisZehua Pei, Hui-Ling Zhen, Lancheng Zou, Xianzhi Yu 等ACL 2026 · 被引用 6 次
- Towards Stable and Effective Reinforcement Learning for Mixture-of-ExpertsDi Zhang, Xun Wu, Shaohan Huang, Lingjie Jiang 等ACL 2026 · 被引用 3 次
- Spark Transformer: Reactivating Sparsity in Transformer FFN and AttentionChong You, Kan Wu, Zhipeng Jia, Lin Chen 等NeurIPS 2025 · 被引用 3 次
它引用的顶会 Paper22
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen 等ICLR 2021 · 被引用 1,954 次
- Mixture-of-Experts with Expert Choice RoutingYanqi Zhou, Tao Lei, Hanxiao Liu, Nan Du 等NeurIPS 2022 · 被引用 933 次
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 被引用 852 次
- Efficient large-scale language model training on GPU clusters using megatron-LMDeepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley 等SC 2021 · 被引用 576 次
相关 Paper
- SoftMoE: Soft Differentiable Routing for Mixture-of-Experts in LLMsMikołaj Zasada, Łukasz Struski, Jacek Tabor, Marcin KurdzielICML 2026 · 被引用 2 次
- DirMoE: Dirichlet-Routed Mixture of ExpertsAmirhossein Vahidi, Hesam Asadollahzadeh, Navid Akhavan Attar, Marie Moullet 等ICLR 2026 · 被引用 2 次
- Synergistic Intra- and Cross-Layer Regularization Losses for MoE Expert SpecializationRizhen Hu, Yuan Cao, Boao Kong, Mou Sun 等ICML 2026
- Tight Clusters Make Specialized ExpertsStefan K. Nielsen, Rachel S. Y. Teo, Laziz U. Abdullaev, Tan Minh NguyenICLR 2025
- DTop-p MoE: Sparsity-Controlled Dynamic Top-p MoE for Foundation Model Pre-trainingCan Jin, Hongwu Peng, Mingcan Xiang, Qixin Zhang 等ICML 2026 · 被引用 3 次
