Statistical Advantages of Perturbing Cosine Router in Mixture of Experts
Huy Nguyen, Pedram Akbarian, Huyen Trang Pham, Thien Trang Nguyen Vu, Shujian Zhang, Nhat Ho
摘要
The cosine router in Mixture of Experts (MoE) has recently emerged as an attractive alternative to the conventional linear router. Indeed, the cosine router demonstrates favorable performance in image and language tasks and exhibits better ability to mitigate the representation collapse issue, which often leads to parameter redundancy and limited representation potentials. Despite its empirical success, a comprehensive analysis of the cosine router in MoE has been lacking. Considering the least square estimation of the cosine routing MoE, we demonstrate that due to the intrinsic interaction of the model parameters in the cosine router via some partial differential equations, regardless of the structures of the experts, the estimation rates of experts and model parameters can be as slow as O(1/ log τ (n)) where τ > 0 is some constant and n is the sample size. Surprisingly, these pessimistic non-polynomial convergence rates can be circumvented by the widely used technique in practice to stabilize the cosine router -simply adding noises to the ℓ 2 -norms in the cosine router, which we refer to as perturbed cosine router. Under the strongly identifiable settings of the expert functions, we prove that the estimation rates for both the experts and model parameters under the perturbed cosine routing MoE are significantly improved to polynomial rates. Finally, we conduct extensive simulation studies in both synthetic and real data settings to empirically validate our theoretical results.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Sigmoid Gating is More Sample Efficient than Softmax Gating in Mixture of ExpertsHuy Nguyen, Nhat Ho, Alessandro RinaldoNeurIPS 2024 · 被引用 35 次
- One-Prompt Strikes Back: Sparse Mixture of Experts for Prompt-based Continual LearningMinh Le, Bao-Ngoc Dao, Huy Nguyen, Quyen Tran 等ICLR 2026 · 被引用 3 次
- On the Expressive Power of Mixture-of-Experts for Structured Complex TasksMingze Wang, Weinan ENeurIPS 2025 · 被引用 3 次
- On Zero-Initialized Attention: Optimal Prompt and Gating Factor EstimationNghiem Tuong Diep, Huy Nguyen, Chau Nguyen, Minh Le 等ICML 2025
- RepLoRA: Reparameterizing Low-rank Adaptation via the Perspective of Mixture of ExpertsTuan Truong, Chau Nguyen, Huy Nguyen, Minh Le 等ICML 2025
它引用的顶会 Paper21
- In Search of Lost Domain GeneralizationIshaan Gulrajani, David Lopez-PazICLR 2021 · 被引用 1,416 次
- Scaling Vision with Sparse Mixture of ExpertsCarlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann 等NeurIPS 2021 · 被引用 1,213 次
- From Sparse to Soft Mixtures of ExpertsJoan Puigcerver, Carlos Riquelme Ruiz, Basil Mustafa, Neil HoulsbyICLR 2024 · 被引用 264 次
- On the Representation Collapse of Sparse Mixture of ExpertsZewen Chi, Li Dong, Shaohan Huang, Damai Dai 等NeurIPS 2022 · 被引用 223 次
- DSelect-k: Differentiable Selection in the Mixture of Experts with Applications to Multi-Task LearningHussein Hazimeh, Zhe Zhao, Aakanksha Chowdhery, Maheswaran Sathiamoorthy 等NeurIPS 2021 · 被引用 216 次
相关 Paper
- Tight Clusters Make Specialized ExpertsStefan K. Nielsen, Rachel S. Y. Teo, Laziz U. Abdullaev, Tan Minh NguyenICLR 2025
- Is Temperature Sample Efficient for Softmax Gaussian Mixture of Experts?Huy Nguyen, Pedram Akbarian, Nhat HoICML 2024 · 被引用 20 次
- On Least Square Estimation in Softmax Gating Mixture of ExpertsHuy Nguyen, Nhat Ho, Alessandro RinaldoICML 2024 · 被引用 25 次
- STAR: Rethinking MoE Routing as Structure-Aware Subspace LearningSumin Park, Noseong ParkICML 2026
- Towards Understanding the Mixture-of-Experts Layer in Deep LearningZixiang Chen, Yihe Deng, Yue Wu, Quanquan Gu 等NeurIPS 2022 · 被引用 199 次
