Learning Mixtures of Experts with EM: A Mirror Descent Perspective
Quentin Fruytier, Aryan Mokhtari, Sujay Sanghavi
摘要
Classical Mixtures of Experts (MoE) are Machine Learning models that involve partitioning the input space, with a separate "expert" model trained on each partition. Recently, MoE-based model architectures have become popular as a means to reduce training and inference costs. There, the partitioning function and the experts are both learnt jointly via gradient descent-type methods on the log-likelihood. In this paper we study theoretical guarantees of the Expectation Maximization (EM) algorithm for the training of MoE models. We first rigorously analyze EM for MoE where the conditional distribution of the target and latent variable conditioned on the feature variable belongs to an exponential family of distributions and show its equivalence to projected Mirror Descent with unit step size and a Kullback-Leibler Divergence regularizer. This perspective allows us to derive new convergence results and identify conditions for local linear convergence; In the special case of mixture of 2 linear or logistic experts, we additionally provide guarantees for linear convergence based on the signal-to-noise ratio. Experiments on synthetic and (small-scale) real-world data supports that EM outperforms the gradient descent algorithm both in terms of convergence rate and the achieved accuracy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Robustness of Mixtures of Experts to Feature NoiseDong Sun, Rahul Nittala, Rebekka BurkholzICML 2026 · 被引用 1 次
- -Balancing for Mixture-of-Experts TrainingLizhang Chen, Jonathan Li, Qi Wang, Runlong Liao 等ICML 2026
它引用的顶会 Paper2
- On the SDEs and Scaling Rules for Adaptive Gradient AlgorithmsSadhika Malladi, Kaifeng Lyu, Abhishek Panigrahi, Sanjeev AroraNeurIPS 2022 · 被引用 125 次
- Expected Information Maximization: Using the I-Projection for Mixture Density EstimationPhilipp Becker, Oleg Arenz, Gerhard NeumannICLR 2020 · 被引用 16 次
相关 Paper
- Mixture of Experts Provably Detect and Learn the Latent Cluster Structure in Gradient-Based LearningRyotaro Kawata, Kohsei Matsutani, Yuri Kinoshita, Naoki Nishikawa 等ICML 2025
- A General Theory for Softmax Gating Multinomial Logistic Mixture of ExpertsHuy Nguyen, Pedram Akbarian, TrungTin Nguyen, Nhat HoICML 2024 · 被引用 28 次
- Agnostic Learning of Mixed Linear Regressions with EM and AM AlgorithmsAvishek Ghosh, Arya MazumdarICML 2024 · 被引用 1 次
- Rethinking Convergence in MoE Training: The Role of Routing SparsityWeihao Zhu, Long Shi, Kang Wei, Zhe Wang 等ICML 2026
- Mirror Descent with Relative Smoothness in Measure Spaces, with application to Sinkhorn and EMPierre-Cyril Aubin-Frankowski, Anna Korba, Flavien LégerNeurIPS 2022 · 被引用 61 次
