On the Representation Collapse of Sparse Mixture of Experts
Zewen Chi, Li Dong, Shaohan Huang, Damai Dai, Shuming Ma, Barun Patra, Saksham Singhal, Payal Bajaj, Xia Song, Xian-Ling Mao, Heyan Huang, Furu Wei
Abstract
Sparse mixture of experts provides larger model capacity while requiring a constant computational overhead. It employs the routing mechanism to distribute input tokens to the best-matched experts according to their hidden representations. However, learning such a routing mechanism encourages token clustering around expert centroids, implying a trend toward representation collapse. In this work, we propose to estimate the routing scores between tokens and experts on a low-dimensional hypersphere. We conduct extensive experiments on cross-lingual language model pre-training and fine-tuning on downstream tasks. Experimental results across seven multilingual benchmarks show that our method achieves consistent gains. We also present a comprehensive analysis on the representation and routing behaviors of our models. Our method alleviates the representation collapse issue and achieves more consistent routing than the baseline mixture-of-experts methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2e35d089-5eb6-4470-9741-afe763ce3ceaCited by top-tier papers63
- Language Is Not All You Need: Aligning Perception with Language ModelsShaohan Huang, Li Dong, Wenhui Wang, Yaru Hao et al.NeurIPS 2023 · 810 citations
- Accelerating Distributed MoE Training and Inference with LinaJiamin Li, Yimin Jiang, Yibo Zhu, Cong Wang et al.USENIX ATC 2023 · 191 citations
- FuseMoE: Mixture-of-Experts Transformers for Fleximodal FusionXing Han, Huy Nguyen, Carl Harris, Nhat Ho et al.NeurIPS 2024 · 129 citations
- AdaMV-MoE: Adaptive Multi-Task Vision Mixture-of-ExpertsTianlong Chen, Xuxi Chen, Xianzhi Du, Abdullah Rashwan et al.ICCV 2023 · 119 citations
- Merge, Then Compress: Demystify Efficient SMoE with Hints from Its Routing PolicyPingzhi Li, Zhenyu Zhang, Prateek Yadav, Yi-Lin Sung et al.ICLR 2024 · 97 citations
Builds on21
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- Scaling Vision with Sparse Mixture of ExpertsCarlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann et al.NeurIPS 2021 · 1,213 citations
- GLaM: Efficient Scaling of Language Models with Mixture-of-ExpertsNan Du, Yanping Huang, Andrew M. Dai, Simon Tong et al.ICML 2022 · 1,173 citations
Related papers
- Breaking the Echo Chamber: A Dynamic Ensemble Pruning Perspective on MoEXinlai Kang, Dunyao Xue, Zhengbo Wang, Chengshuo Du et al.ICML 2026
- Multilingual Routing in Mixture-of-ExpertsLucas Bandarkar, Chenyuan Yang, Mohsen Fayyaz, Junlin Hu et al.ICLR 2026 · 34 citations
- Hierarchical Mixture of Experts with Two-Stage OptimizationGleb Molodtsov, Alexander Miasnikov, Aleksandr BeznosikovKDD 2026 · 2 citations
- Expert Divergence Learning for MoE-based Language ModelsJiaang Li, Haibin Chen, Langming Liu, Yujin Yuan et al.ICLR 2026 · 3 citations
- Routing Manifold Alignment Improves Generalization of Mixture-of-Experts LLMsZhongyang Li, Ziyue Li, Tianyi ZhouICLR 2026 · 5 citations
