A General Theory for Softmax Gating Multinomial Logistic Mixture of Experts
Huy Nguyen, Pedram Akbarian, TrungTin Nguyen, Nhat Ho
摘要
Mixture-of-experts (MoE) model incorporates the power of multiple submodels via gating functions to achieve greater performance in numerous regression and classification applications. From a theoretical perspective, while there have been previous attempts to comprehend the behavior of that model under the regression settings through the convergence analysis of maximum likelihood estimation in the Gaussian MoE model, such analysis under the setting of a classification problem has remained missing in the literature. We close this gap by establishing the convergence rates of density estimation and parameter estimation in the softmax gating multinomial logistic MoE model. Notably, when part of the expert parameters vanish, these rates are shown to be slower than polynomial rates owing to an inherent interaction between the softmax gating and expert functions via partial differential equations. To address this issue, we propose using a novel class of modified softmax gating functions which transform the input before delivering them to the gating functions. As a result, the previous interaction disappears and the parameter estimation rates are significantly improved.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- FuseMoE: Mixture-of-Experts Transformers for Fleximodal FusionXing Han, Huy Nguyen, Carl Harris, Nhat Ho 等NeurIPS 2024 · 被引用 129 次
- Sigmoid Gating is More Sample Efficient than Softmax Gating in Mixture of ExpertsHuy Nguyen, Nhat Ho, Alessandro RinaldoNeurIPS 2024 · 被引用 35 次
- On Least Square Estimation in Softmax Gating Mixture of ExpertsHuy Nguyen, Nhat Ho, Alessandro RinaldoICML 2024 · 被引用 25 次
- Is Temperature Sample Efficient for Softmax Gaussian Mixture of Experts?Huy Nguyen, Pedram Akbarian, Nhat HoICML 2024 · 被引用 20 次
- On Minimax Estimation of Parameters in Softmax-Contaminated Mixture of ExpertsFanqi Yan, Huy Nguyen, Dung Le, Pedram Akbarian 等NeurIPS 2025 · 被引用 1 次
它引用的顶会 Paper5
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Mixture-of-Experts with Expert Choice RoutingYanqi Zhou, Tao Lei, Hanxiao Liu, Nan Du 等NeurIPS 2022 · 被引用 933 次
- DSelect-k: Differentiable Selection in the Mixture of Experts with Applications to Multi-Task LearningHussein Hazimeh, Zhe Zhao, Aakanksha Chowdhery, Maheswaran Sathiamoorthy 等NeurIPS 2021 · 被引用 216 次
- A Mixture-of-Expert Approach to RL-based Dialogue ManagementYinlam Chow, Aza Tulepbergenov, Ofir Nachum, Dhawal Gupta 等ICLR 2023 · 被引用 2 次
- Mod-Squad: Designing Mixtures of Experts As Modular Multi-Task LearnersZitian Chen, Yikang Shen, Mingyu Ding, Zhenfang Chen 等CVPR 2023
相关 Paper
- Demystifying Softmax Gating Function in Gaussian Mixture of ExpertsHuy Nguyen, TrungTin Nguyen, Nhat HoNeurIPS 2023 · 被引用 44 次
- Statistical Perspective of Top-K Sparse Softmax Gating Mixture of ExpertsHuy Nguyen, Pedram Akbarian, Fanqi Yan, Nhat HoICLR 2024 · 被引用 29 次
- Theory on Mixture-of-Experts in Continual LearningHongbo Li, Sen Lin, Lingjie Duan, Yingbin Liang 等ICLR 2025
- Learning Mixtures of Experts with EM: A Mirror Descent PerspectiveQuentin Fruytier, Aryan Mokhtari, Sujay SanghaviICML 2025
- Statistical Advantages of Perturbing Cosine Router in Mixture of ExpertsHuy Nguyen, Pedram Akbarian, Huyen Trang Pham, Thien Trang Nguyen Vu 等ICLR 2025
