MomentumSMoE: Integrating Momentum into Sparse Mixture of Experts
Rachel S. Y. Teo, Tan M. Nguyen
摘要
Sparse Mixture of Experts (SMoE) has become the key to unlocking unparalleled scalability in deep learning. SMoE has the potential to exponentially increase parameter count while maintaining the efficiency of the model by only activating a small subset of these parameters for a given sample. However, it has been observed that SMoE suffers from unstable training and has difficulty adapting to new distributions, leading to the model's lack of robustness to data contamination. To overcome these limitations, we first establish a connection between the dynamics of the expert representations in SMoEs and gradient descent on a multi-objective optimization problem. Leveraging our framework, we then integrate momentum into SMoE and propose a new family of SMoEs named MomentumSMoE. We theoretically prove and numerically demonstrate that MomentumSMoE is more stable and robust than SMoE. In particular, we verify the advantages of MomentumSMoE over SMoE on a variety of practical tasks including ImageNet-1K object recognition and WikiText-103 language modeling. We demonstrate the applicability of MomentumSMoE to many types of SMoE models, including those in the Sparse MoE model for vision (V-MoE) and the Generalist Language Model (GLaM). We also show that other advanced momentum-based optimization methods, such as Adam, can be easily incorporated into the MomentumSMoE framework for designing new SMoE models with even better performance, almost negligible additional computation cost, and simple implementations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- OneSparse: A Unified Framework for Sparse Activation Layers in Vision ModelsXingkui Zhu, Dingkang Liang, Cheng Chen, Daoxin Zhang 等CVPR 2026
- MoLEx: Mixture of Layer Experts for Fine-tuning with Sparse UpcyclingRachel S. Y. Teo, Tan Minh NguyenICLR 2025
- Tight Clusters Make Specialized ExpertsStefan K. Nielsen, Rachel S. Y. Teo, Laziz U. Abdullaev, Tan Minh NguyenICLR 2025
它引用的顶会 Paper35
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
相关 Paper
- Input Domain Aware MoE: Decoupling Routing Decisions from Task Optimization in Mixture of ExpertsYongXiang Hua, Haoyu Cao, Zhou Tao, Bocheng Li 等ACM MM 2025 · 被引用 1 次
- Multi-Head Mixture-of-ExpertsXun Wu, Shaohan Huang, Wenhui Wang, Shuming Ma 等NeurIPS 2024 · 被引用 42 次
- Mixture of Tokens: Continuous MoE through Cross-Example AggregationSzymon Antoniak, Michal Krutul, Maciej Pióro, Jakub Krajewski 等NeurIPS 2024 · 被引用 6 次
- GMoE: Global Mixture of Experts with Logit PropagationGeonwoo Hong, Taehwan KimACL 2026
- Expert Merging in Sparse Mixture of Experts with Nash BargainingDung Viet Nguyen, Anh Nguyen Thi, Minh Hoang Nguyen, Luc Nguyen 等ICLR 2026 · 被引用 3 次
