Taming Sparsely Activated Transformer with Stochastic Experts
Simiao Zuo, Xiaodong Liu, Jian Jiao, Young Jin Kim, Hany Hassan, Ruofei Zhang, Jianfeng Gao, Tuo Zhao
Abstract
Sparsely activated models (SAMs), such as Mixture-of-Experts (MoE), can easily scale to have outrageously large amounts of parameters without significant increase in computational cost. However, SAMs are reported to be parameter inefficient such that larger models do not always lead to better performance. While most on-going research focuses on improving SAMs models by exploring methods of routing inputs to experts, our analysis reveals that such research might not lead to the solution we expect, i.e., the commonly-used routing methods based on gating mechanisms do not work better than randomly routing inputs to experts. In this paper, we propose a new expert-based model, THOR (Transformer witH StOchastic ExpeRts). Unlike classic expert-based models, such as the Switch Transformer (Fedus et al., 2021) , experts in THOR are randomly activated for each input during training and inference. THOR models are trained using a consistency regularized loss, where experts learn not only from training data but also from other experts as teachers, such that all the experts make consistent predictions. We validate the effectiveness of THOR on machine translation tasks. Results show that THOR models are more parameter efficient in that they significantly outperform the Transformer and MoE models across various settings. For example, in multilingual translation, THOR outperforms the Switch Transformer by 2 BLEU scores, and obtains the same BLEU score as that of a state-of-the-art MoE model (Kim et al., 2021) that is 18 times larger. Our code is publicly available at: https://github.com/microsoft/ Stochastic-Mixture-of-Experts .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f11c2985-293d-4786-bddb-eaa515cddcdbCited by top-tier papers69
- DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI ScaleSamyam Rajbhandari, Conglong Li, Zhewei Yao, Minjia Zhang et al.ICML 2022 · 523 citations
- Accelerating Distributed MoE Training and Inference with LinaJiamin Li, Yimin Jiang, Yibo Zhu, Cong Wang et al.USENIX ATC 2023 · 191 citations
- Pushing Mixture of Experts to the Limit: Extremely Parameter Efficient MoE for Instruction TuningTed Zadouri, Ahmet Üstün, Arash Ahmadian, Beyza Ermis et al.ICLR 2024 · 169 citations
- AdaMV-MoE: Adaptive Multi-Task Vision Mixture-of-ExpertsTianlong Chen, Xuxi Chen, Xianzhi Du, Abdullah Rashwan et al.ICCV 2023 · 119 citations
- Flex-MoE: Modeling Arbitrary Modality Combination via the Flexible Mixture-of-ExpertsSukwon Yun, Inyoung Choi, Jie Peng, Yangfan Wu et al.NeurIPS 2024 · 98 citations
Builds on11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
- On the Variance of the Adaptive Learning Rate and BeyondLiyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen et al.ICLR 2020 · 2,210 citations
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
Related papers
- Gating Dropout: Communication-efficient Regularization for Sparsely Activated TransformersRui Liu, Young Jin Kim, Alexandre Muzio, Hany HassanICML 2022 · 31 citations
- THOR-MoE: Hierarchical Task-Guided and Context-Responsive Routing for Neural Machine TranslationYunlong Liang, Fandong Meng, Jie ZhouACL 2025 · 1 citation
- Exploiting Activation Sparsity with Dense to Dynamic-k Mixture-of-Experts ConversionFilip Szatkowski, Bartosz Wójcik, Mikolaj Piórczynski, Simone ScardapaneNeurIPS 2024 · 19 citations
- ResMoE: Space-efficient Compression of Mixture of Experts LLMs via Residual RestorationMengting Ai, Tianxin Wei, Yifan Chen, Zhichen Zeng et al.KDD 2025 · 5 citations
- StableMoE: Stable Routing Strategy for Mixture of ExpertsDamai Dai, Li Dong, Shuming Ma, Bo Zheng et al.ACL 2022
