Adaptive Gating in Mixture-of-Experts based Language Models
Jiamin Li, Qiang Su, Yitao Yang, Yimin Jiang, Cong Wang, Hong Xu
摘要
Large language models have demonstrated exceptional language understanding capabilities in many NLP tasks. Sparsely activated mixture-of-experts (MoE) has emerged as a promising solution for scaling models while maintaining a constant number of computational operations. Existing MoE models adopt a fixed gating network where each token is computed by the same number of experts. This contradicts our intuition that the tokens in each sequence vary in terms of their linguistic complexity and, consequently, require different computational costs. Little is discussed in prior research on the trade-off between computation per token and model performance. This paper introduces adaptive gating in MoE, a flexible training strategy that allows tokens to be processed by a variable number of experts based on expert probability distribution. Adaptive gating preserves sparsity while improving training efficiency. We further draw upon curriculum learning to better align the order of training samples and maximize the training time savings. Extensive experiments on diverse NLP tasks show that adaptive gating reduces at most 22.5% training time while maintaining inference quality. Moreover, we conduct a comprehensive analysis of the gating decisions and present our insights on which tokens are inherently difficult to process, depending on the specific language task.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- HybriMoE: Hybrid CPU-GPU Scheduling and Cache Management for Efficient MoE InferenceShuzhang Zhong, Yanfan Sun, Ling Liang, Runsheng Wang 等DAC 2025 · 被引用 8 次
- THOR-MoE: Hierarchical Task-Guided and Context-Responsive Routing for Neural Machine TranslationYunlong Liang, Fandong Meng, Jie ZhouACL 2025 · 被引用 1 次
- Cooperation of Experts: Fusing Heterogeneous Information with Large MarginShuo Wang, Shunyang Huang, Jinghui Yuan, Zhixiang Shen 等ICML 2025
- FIRM-MoE: Fine-GrainedExpert Decomposition for Resource-Adaptive MoE InferenceKeyu Chen, Qihang Zhou, Bin Qian, Zhenyu Wen 等AAAI 2026
它引用的顶会 Paper8
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen 等ICLR 2021 · 被引用 1,954 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- GLaM: Efficient Scaling of Language Models with Mixture-of-ExpertsNan Du, Yanping Huang, Andrew M. Dai, Simon Tong 等ICML 2022 · 被引用 1,173 次
- Mixture-of-Experts with Expert Choice RoutingYanqi Zhou, Tao Lei, Hanxiao Liu, Nan Du 等NeurIPS 2022 · 被引用 933 次
- DSelect-k: Differentiable Selection in the Mixture of Experts with Applications to Multi-Task LearningHussein Hazimeh, Zhe Zhao, Aakanksha Chowdhery, Maheswaran Sathiamoorthy 等NeurIPS 2021 · 被引用 216 次
相关 Paper
- Dynamic Mixture of Experts: An Auto-Tuning Approach for Efficient Transformer ModelsYongxin Guo, Zhenglin Cheng, Xiaoying Tang, Zhaopeng Tu 等ICLR 2025
- PipeMoE: Accelerating Mixture-of-Experts through Adaptive PipeliningShaohuai Shi, Xinglin Pan, Xiaowen Chu, Bo LiINFOCOM 2023 · 被引用 23 次
- FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts ModelsXinglin Pan, Wenxiang Lin, Lin Zhang, Shaohuai Shi 等ASPLOS 2025 · 被引用 12 次
- Gating Dropout: Communication-efficient Regularization for Sparsely Activated TransformersRui Liu, Young Jin Kim, Alexandre Muzio, Hany HassanICML 2022 · 被引用 31 次
- DSMoE: Matrix-Partitioned Experts with Dynamic Routing for Computation-Efficient Dense LLMsMinxuan Lv, Zhenpeng Su, Leiyu Pan, Yizhe Xiong 等EMNLP 2025
