MoEC: Mixture of Expert Clusters
Yuan Xie, Shaohan Huang, Tianyu Chen, Furu Wei
摘要
Sparsely Mixture of Experts (MoE) has received great interest due to its promising scaling capability with affordable computational overhead. MoE models convert dense layers into sparse experts, and utilize a gated routing network to make experts conditionally activated. However, as the number of experts grows, MoE with outrageous parameters suffers from overfitting and sparse data allocation. Such problems are especially severe on tasks with limited data, thus hindering the progress towards improving performance by scaling up. We verify that there exists a performance upper bound of scaling up sparse MoE. In this work, we propose Mixture of Expert Clusters — a general approach to enable expert layers to learn more diverse and appropriate knowledge by imposing variance-based constraints on the routing stage. Given this, we could further propose a cluster-level expert dropout strategy specifically designed for the expert cluster structure. Our experiments reveal that MoEC could improve performance on machine translation and natural language understanding tasks. MoEC plays a positive role in mitigating overfitting and sparse data allocation problems, thus fully releasing the potential of large-scale sparse models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Mixture of LoRA ExpertsXun Wu, Shaohan Huang, Furu WeiICLR 2024 · 被引用 174 次
- Multi-Head Mixture-of-ExpertsXun Wu, Shaohan Huang, Wenhui Wang, Shuming Ma 等NeurIPS 2024 · 被引用 42 次
- UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context LearningZihao Huang, Yu Bao, Qiyang Min, Siyan Chen 等ICLR 2026 · 被引用 6 次
- On Token's Dilemma: Dynamic MoE with Drift-Aware Token Assignment for Continual Learning of Large Vision Language ModelsChongyang Zhao, Mingsong Li, Haodong Lu, Dong GongCVPR 2026 · 被引用 3 次
- Effective Node-Level Anomaly Detection in HPC Systems via Coarse-Grained Clustering and Fine-Grained Model SharingSibo Xia, Yongqian Sun, Xijie Pan, Yuan Yuan 等SC 2025 · 被引用 3 次
它引用的顶会 Paper10
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 被引用 3,632 次
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen 等ICLR 2021 · 被引用 1,954 次
- Scaling Vision with Sparse Mixture of ExpertsCarlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann 等NeurIPS 2021 · 被引用 1,213 次
- BASE Layers: Simplifying Training of Large, Sparse ModelsMike Lewis, Shruti Bhosale, Tim Dettmers, Naman Goyal 等ICML 2021 · 被引用 382 次
相关 Paper
- Gating Dropout: Communication-efficient Regularization for Sparsely Activated TransformersRui Liu, Young Jin Kim, Alexandre Muzio, Hany HassanICML 2022 · 被引用 31 次
- Tight Clusters Make Specialized ExpertsStefan K. Nielsen, Rachel S. Y. Teo, Laziz U. Abdullaev, Tan Minh NguyenICLR 2025
- Mixture-of-Clustered-Experts: Advancing Expert Specialization and Generalization in Instruction TuningSugyeong Eo, Jung Jun Lee, Chanjun Park, Heuiseok LimEMNLP 2025
- Towards Understanding the Mixture-of-Experts Layer in Deep LearningZixiang Chen, Yihe Deng, Yue Wu, Quanquan Gu 等NeurIPS 2022 · 被引用 199 次
- Adaptive Gating in Mixture-of-Experts based Language ModelsJiamin Li, Qiang Su, Yitao Yang, Yimin Jiang 等EMNLP 2023 · 被引用 9 次
