ProbMoE: Differentiable Probabilistic Routing for Mixture-of-Experts
Heng Zhao, Zilei Shao, Guy Van den Broeck, Zhe Zeng
摘要
Mixture-of-Experts (MoE) models scale by activating only a small subset of experts per token. However, training such models remains challenging because top- routing is discrete and non-differentiable, requiring gradient estimators for expert selection whose design remains a central open problem. We introduce ProbMoE, a probabilistic routing framework that models expert selection as a distribution over cardinality-constrained expert subsets and formulates routing as probabilistic inference in this discrete subset space. We first propose ProbMoE Exact- routing, which samples -expert subsets in the forward pass, and the backward pass uses gradients through each expert's exact marginal probability as a tractable surrogate for the true gradient. ProbMoE naturally generalizes to a dynamic- routing setting, where both training and inference constrain the routing cardinality to the same predefined range, allowing adaptive expert allocation per token. Across benchmarks and model backbones, ProbMoE Exact- achieves strong performance compared to competitive baselines, with improved expert utilization and routing diversity; ProbMoE Dynamic- achieves comparable performance with fewer activated experts. Code is available at: https://github.com/HengHugoZhao/ProbMoE.git
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- GLaM: Efficient Scaling of Language Models with Mixture-of-ExpertsNan Du, Yanping Huang, Andrew M. Dai, Simon Tong 等ICML 2022 · 被引用 1,173 次
- DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI ScaleSamyam Rajbhandari, Conglong Li, Zhewei Yao, Minjia Zhang 等ICML 2022 · 被引用 523 次
- BASE Layers: Simplifying Training of Large, Sparse ModelsMike Lewis, Shruti Bhosale, Tim Dettmers, Naman Goyal 等ICML 2021 · 被引用 382 次
- Unified Scaling Laws for Routed Language ModelsAidan Clark, Diego de Las Casas, Aurelia Guy, Arthur Mensch 等ICML 2022 · 被引用 266 次
相关 Paper
- ReMoE: Fully Differentiable Mixture-of-Experts with ReLU RoutingZiteng Wang, Jun Zhu, Jianfei ChenICLR 2025
- SoftMoE: Soft Differentiable Routing for Mixture-of-Experts in LLMsMikołaj Zasada, Łukasz Struski, Jacek Tabor, Marcin KurdzielICML 2026 · 被引用 2 次
- Dense Backpropagation Improves Training for Sparse Mixture-of-ExpertsAshwinee Panda, Vatsal Baherwani, Zain Sarwar, Benjamin Thérien 等NeurIPS 2025 · 被引用 10 次
- DirMoE: Dirichlet-Routed Mixture of ExpertsAmirhossein Vahidi, Hesam Asadollahzadeh, Navid Akhavan Attar, Marie Moullet 等ICLR 2026 · 被引用 2 次
- DTop-p MoE: Sparsity-Controlled Dynamic Top-p MoE for Foundation Model Pre-trainingCan Jin, Hongwu Peng, Mingcan Xiang, Qixin Zhang 等ICML 2026 · 被引用 3 次
