On the Benefits of Learning to Route in Mixture-of-Experts Models
Nishanth Dikkala, Nikhil Ghosh, Raghu Meka, Rina Panigrahy, Nikhil Vyas, Xin Wang
摘要
Mixture-of-Expert (MoE) Transformer models, such as the Switch Transformer, allow us to successfully scale up model sizes while keeping the amount of compute time fixed. Prior work has established the computational efficiency benefits of using these models. A core component of these models is a router that routes input tokens to different experts in a layer. We show theoretical and empirical evidence that the router's ability to route tokens intelligently confers a significant advantage to MoE models. We study synthetic settings where the input data is distributed in clusters and show theoretically and empirically that the router learns to route the inputs according to these clusters. Then we perform experiments on real data using the T5X library, where we observe that a trainable router confers a non-trivial benefit instead of a non-trainable router.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- On the Expressive Power of Mixture-of-Experts for Structured Complex TasksMingze Wang, Weinan ENeurIPS 2025 · 被引用 3 次
- ReMoE: Boosting Expert Reuse through Router Fine-Tuning in Memory-Constrained MoE LLM InferenceXiongwei Zhu, Xiaojian Liao, Tianyang Jiang, Yusen Zhang 等ICML 2026 · 被引用 2 次
- Enhancing Mixture-of-Experts Specialization via Cluster-Aware UpcyclingSanghyeok Chu, Pyunghwan Ahn, Gwangmo Song, Seung Hwan Kim 等CVPR 2026
- Tight Clusters Make Specialized ExpertsStefan K. Nielsen, Rachel S. Y. Teo, Laziz U. Abdullaev, Tan Minh NguyenICLR 2025
它引用的顶会 Paper11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen 等ICLR 2021 · 被引用 1,954 次
- Scaling Vision with Sparse Mixture of ExpertsCarlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann 等NeurIPS 2021 · 被引用 1,213 次
- GLaM: Efficient Scaling of Language Models with Mixture-of-ExpertsNan Du, Yanping Huang, Andrew M. Dai, Simon Tong 等ICML 2022 · 被引用 1,173 次
- Mixture-of-Experts with Expert Choice RoutingYanqi Zhou, Tao Lei, Hanxiao Liu, Nan Du 等NeurIPS 2022 · 被引用 933 次
相关 Paper
- StableMoE: Stable Routing Strategy for Mixture of ExpertsDamai Dai, Li Dong, Shuming Ma, Bo Zheng 等ACL 2022
- Union-of-Experts: Neurons in Mixture-of-Experts are Secretly RoutersSonghao Wu, Ang Lv, Ruobing Xie, Samm Sun 等ACL 2026
- Layerwise Recurrent Router for Mixture-of-ExpertsZihan Qiu, Zeyu Huang, Shuang Cheng, Yizhi Zhou 等ICLR 2025
- Hyperparameter Transfer with Mixture-of-Expert LayersTianze Jiang, Blake Bordelon, Cengiz Pehlevan, Boris HaninICML 2026 · 被引用 6 次
- Harder Task Needs More Experts: Dynamic Routing in MoE ModelsQuzhe Huang, Zhenwei An, Nan Zhuang, Mingxu Tao 等ACL 2024 · 被引用 11 次
