Rethinking Sparse Mixture of Experts from a Unified Perspective
Giang Do, Hung Le, Truyen Tran
摘要
Sparse Mixture of Experts (SMoE) models scale the capacity of models while maintaining constant computational overhead. SMoE methods fall into two categories: Token Choice, which routes each token to a fixed number of experts, and Expert Choice, which assigns a fixed number of tokens to each expert. However, the use of fixed budgets for tokens or experts causes both approaches to select irrelevant token-expert pairs or overlook critical assignments, which degrades overall performance. To fill that gap, we rethink SMoE from a unified perspective through the lens of linear programming, which provides a general formulation for SMoE models. Furthermore, we introduce Unified Sparse Mixture of Experts (USMoE), a novel framework comprising a unified mechanism and a unified score to overcome these limitations. We provide both theoretical justification and empirical evidence demonstrating USMoE's effectiveness. Extensive evaluations across diverse data settings (clean and corrupted), multiple domains (including texts and vision tasks), and different learning approaches (training-free and training-based) show that USMoE not only delivers significant performance improvements over existing SMoE methods, but also enables more flexible expert selection budgets, reducing inference costs without compromising model performance. Our implementation is publicly available at https://github.com/giangdip241 0/USMoE .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper17
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
- Scaling Vision with Sparse Mixture of ExpertsCarlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann 等NeurIPS 2021 · 被引用 1,213 次
- GaLore: Memory-Efficient LLM Training by Gradient Low-Rank ProjectionJiawei Zhao, Zhenyu Zhang, Beidi Chen, Zhangyang Wang 等ICML 2024 · 被引用 433 次
相关 Paper
- Eigenvectors of Experts are Training-free Non-collapsing RoutersGiang Do, Hung Le, Truyen TranICML 2026 · 被引用 1 次
- Input Domain Aware MoE: Decoupling Routing Decisions from Task Optimization in Mixture of ExpertsYongXiang Hua, Haoyu Cao, Zhou Tao, Bocheng Li 等ACM MM 2025 · 被引用 1 次
- Multi-Head Mixture-of-ExpertsXun Wu, Shaohan Huang, Wenhui Wang, Shuming Ma 等NeurIPS 2024 · 被引用 42 次
- Dynamic Mixture of Experts: An Auto-Tuning Approach for Efficient Transformer ModelsYongxin Guo, Zhenglin Cheng, Xiaoying Tang, Zhaopeng Tu 等ICLR 2025
- HMoE: Heterogeneous Mixture of Experts for Language ModelingAn Wang, Xingwu Sun, Ruobing Xie, Shuaipeng Li 等EMNLP 2025 · 被引用 2 次
