Is Temperature Sample Efficient for Softmax Gaussian Mixture of Experts?
Huy Nguyen, Pedram Akbarian, Nhat Ho
Abstract
Dense-to-sparse gating mixture of experts (MoE) has recently become an effective alternative to a well-known sparse MoE. Rather than fixing the number of activated experts as in the latter model, which could limit the investigation of potential experts, the former model utilizes the temperature to control the softmax weight distribution and the sparsity of the MoE during training in order to stabilize the expert specialization. Nevertheless, while there are previous attempts to theoretically comprehend the sparse MoE, a comprehensive analysis of the dense-to-sparse gating MoE has remained elusive. Therefore, we aim to explore the impacts of the dense-to-sparse gate on the maximum likelihood estimation under the Gaussian MoE in this paper. We demonstrate that due to interactions between the temperature and other model parameters via some partial differential equations, the convergence rates of parameter estimations are slower than any polynomial rates, and could be as slow as , where denotes the sample size. To address this issue, we propose using a novel activation dense-to-sparse gate, which routes the output of a linear layer to an activation function before delivering them to the softmax function. By imposing linearly independence conditions on the activation function and its derivatives, we show that the parameter estimation rates are significantly improved to polynomial rates. Finally, we conduct a simulation study to empirically validate our theoretical results.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ecb4134e-9e18-41c5-9ba6-24eddacccc5aCited by top-tier papers9
- FuseMoE: Mixture-of-Experts Transformers for Fleximodal FusionXing Han, Huy Nguyen, Carl Harris, Nhat Ho et al.NeurIPS 2024 · 129 citations
- Mixture of Experts Meets Prompt-Based Continual LearningMinh Le, An Nguyen The, Huy Nguyen, Trang Nguyen et al.NeurIPS 2024 · 57 citations
- Sigmoid Gating is More Sample Efficient than Softmax Gating in Mixture of ExpertsHuy Nguyen, Nhat Ho, Alessandro RinaldoNeurIPS 2024 · 35 citations
- Revisit Visual Prompt Tuning: The Expressiveness of Prompt ExpertsMinh Le, Anh Nguyen, Huy Nguyen, Chau Nguyen et al.ICLR 2026 · 6 citations
- Uncertainty-driven Embedding ConvolutionSungjun Lim, Kangjun Noh, Youngjun Choi, Heeyoung Lee et al.ICLR 2026
Builds on7
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- DSelect-k: Differentiable Selection in the Mixture of Experts with Applications to Multi-Task LearningHussein Hazimeh, Zhe Zhao, Aakanksha Chowdhery, Maheswaran Sathiamoorthy et al.NeurIPS 2021 · 216 citations
- Towards Understanding the Mixture-of-Experts Layer in Deep LearningZixiang Chen, Yihe Deng, Yue Wu, Quanquan Gu et al.NeurIPS 2022 · 199 citations
- Demystifying Softmax Gating Function in Gaussian Mixture of ExpertsHuy Nguyen, TrungTin Nguyen, Nhat HoNeurIPS 2023 · 44 citations
- Statistical Perspective of Top-K Sparse Softmax Gating Mixture of ExpertsHuy Nguyen, Pedram Akbarian, Fanqi Yan, Nhat HoICLR 2024 · 29 citations
Related papers
- A General Theory for Softmax Gating Multinomial Logistic Mixture of ExpertsHuy Nguyen, Pedram Akbarian, TrungTin Nguyen, Nhat HoICML 2024 · 28 citations
- On Least Square Estimation in Softmax Gating Mixture of ExpertsHuy Nguyen, Nhat Ho, Alessandro RinaldoICML 2024 · 25 citations
- On Minimax Estimation of Parameters in Softmax-Contaminated Mixture of ExpertsFanqi Yan, Huy Nguyen, Dung Le, Pedram Akbarian et al.NeurIPS 2025 · 1 citation
- MoEC: Mixture of Expert ClustersYuan Xie, Shaohan Huang, Tianyu Chen, Furu WeiAAAI 2023 · 27 citations
- Statistical Advantages of Perturbing Cosine Router in Mixture of ExpertsHuy Nguyen, Pedram Akbarian, Huyen Trang Pham, Thien Trang Nguyen Vu et al.ICLR 2025
