Statistical Advantages of Perturbing Cosine Router in Mixture of Experts
Huy Nguyen, Pedram Akbarian, Huyen Trang Pham, Thien Trang Nguyen Vu, Shujian Zhang, Nhat Ho
Abstract
The cosine router in Mixture of Experts (MoE) has recently emerged as an attractive alternative to the conventional linear router. Indeed, the cosine router demonstrates favorable performance in image and language tasks and exhibits better ability to mitigate the representation collapse issue, which often leads to parameter redundancy and limited representation potentials. Despite its empirical success, a comprehensive analysis of the cosine router in MoE has been lacking. Considering the least square estimation of the cosine routing MoE, we demonstrate that due to the intrinsic interaction of the model parameters in the cosine router via some partial differential equations, regardless of the structures of the experts, the estimation rates of experts and model parameters can be as slow as O(1/ log τ (n)) where τ > 0 is some constant and n is the sample size. Surprisingly, these pessimistic non-polynomial convergence rates can be circumvented by the widely used technique in practice to stabilize the cosine router -simply adding noises to the ℓ 2 -norms in the cosine router, which we refer to as perturbed cosine router. Under the strongly identifiable settings of the expert functions, we prove that the estimation rates for both the experts and model parameters under the perturbed cosine routing MoE are significantly improved to polynomial rates. Finally, we conduct extensive simulation studies in both synthetic and real data settings to empirically validate our theoretical results.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c8bc467f-9017-46c1-bfdb-0d48d2240877Cited by top-tier papers6
- Sigmoid Gating is More Sample Efficient than Softmax Gating in Mixture of ExpertsHuy Nguyen, Nhat Ho, Alessandro RinaldoNeurIPS 2024 · 35 citations
- One-Prompt Strikes Back: Sparse Mixture of Experts for Prompt-based Continual LearningMinh Le, Bao-Ngoc Dao, Huy Nguyen, Quyen Tran et al.ICLR 2026 · 3 citations
- On the Expressive Power of Mixture-of-Experts for Structured Complex TasksMingze Wang, Weinan ENeurIPS 2025 · 3 citations
- On Zero-Initialized Attention: Optimal Prompt and Gating Factor EstimationNghiem Tuong Diep, Huy Nguyen, Chau Nguyen, Minh Le et al.ICML 2025
- RepLoRA: Reparameterizing Low-rank Adaptation via the Perspective of Mixture of ExpertsTuan Truong, Chau Nguyen, Huy Nguyen, Minh Le et al.ICML 2025
Builds on21
- In Search of Lost Domain GeneralizationIshaan Gulrajani, David Lopez-PazICLR 2021 · 1,416 citations
- Scaling Vision with Sparse Mixture of ExpertsCarlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann et al.NeurIPS 2021 · 1,213 citations
- From Sparse to Soft Mixtures of ExpertsJoan Puigcerver, Carlos Riquelme Ruiz, Basil Mustafa, Neil HoulsbyICLR 2024 · 264 citations
- On the Representation Collapse of Sparse Mixture of ExpertsZewen Chi, Li Dong, Shaohan Huang, Damai Dai et al.NeurIPS 2022 · 223 citations
- DSelect-k: Differentiable Selection in the Mixture of Experts with Applications to Multi-Task LearningHussein Hazimeh, Zhe Zhao, Aakanksha Chowdhery, Maheswaran Sathiamoorthy et al.NeurIPS 2021 · 216 citations
Related papers
- Tight Clusters Make Specialized ExpertsStefan K. Nielsen, Rachel S. Y. Teo, Laziz U. Abdullaev, Tan Minh NguyenICLR 2025
- Is Temperature Sample Efficient for Softmax Gaussian Mixture of Experts?Huy Nguyen, Pedram Akbarian, Nhat HoICML 2024 · 20 citations
- On Least Square Estimation in Softmax Gating Mixture of ExpertsHuy Nguyen, Nhat Ho, Alessandro RinaldoICML 2024 · 25 citations
- STAR: Rethinking MoE Routing as Structure-Aware Subspace LearningSumin Park, Noseong ParkICML 2026
- Towards Understanding the Mixture-of-Experts Layer in Deep LearningZixiang Chen, Yihe Deng, Yue Wu, Quanquan Gu et al.NeurIPS 2022 · 199 citations
