Learning Mixtures of Experts with EM: A Mirror Descent Perspective
Quentin Fruytier, Aryan Mokhtari, Sujay Sanghavi
Abstract
Classical Mixtures of Experts (MoE) are Machine Learning models that involve partitioning the input space, with a separate "expert" model trained on each partition. Recently, MoE-based model architectures have become popular as a means to reduce training and inference costs. There, the partitioning function and the experts are both learnt jointly via gradient descent-type methods on the log-likelihood. In this paper we study theoretical guarantees of the Expectation Maximization (EM) algorithm for the training of MoE models. We first rigorously analyze EM for MoE where the conditional distribution of the target and latent variable conditioned on the feature variable belongs to an exponential family of distributions and show its equivalence to projected Mirror Descent with unit step size and a Kullback-Leibler Divergence regularizer. This perspective allows us to derive new convergence results and identify conditions for local linear convergence; In the special case of mixture of 2 linear or logistic experts, we additionally provide guarantees for linear convergence based on the signal-to-noise ratio. Experiments on synthetic and (small-scale) real-world data supports that EM outperforms the gradient descent algorithm both in terms of convergence rate and the achieved accuracy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1eec8e1c-0446-44e4-92a2-77ca5e7c8cb1Cited by top-tier papers2
- Robustness of Mixtures of Experts to Feature NoiseDong Sun, Rahul Nittala, Rebekka BurkholzICML 2026 · 1 citation
- -Balancing for Mixture-of-Experts TrainingLizhang Chen, Jonathan Li, Qi Wang, Runlong Liao et al.ICML 2026
Builds on2
- On the SDEs and Scaling Rules for Adaptive Gradient AlgorithmsSadhika Malladi, Kaifeng Lyu, Abhishek Panigrahi, Sanjeev AroraNeurIPS 2022 · 125 citations
- Expected Information Maximization: Using the I-Projection for Mixture Density EstimationPhilipp Becker, Oleg Arenz, Gerhard NeumannICLR 2020 · 16 citations
Related papers
- Mixture of Experts Provably Detect and Learn the Latent Cluster Structure in Gradient-Based LearningRyotaro Kawata, Kohsei Matsutani, Yuri Kinoshita, Naoki Nishikawa et al.ICML 2025
- A General Theory for Softmax Gating Multinomial Logistic Mixture of ExpertsHuy Nguyen, Pedram Akbarian, TrungTin Nguyen, Nhat HoICML 2024 · 28 citations
- Agnostic Learning of Mixed Linear Regressions with EM and AM AlgorithmsAvishek Ghosh, Arya MazumdarICML 2024 · 1 citation
- Rethinking Convergence in MoE Training: The Role of Routing SparsityWeihao Zhu, Long Shi, Kang Wei, Zhe Wang et al.ICML 2026
- Mirror Descent with Relative Smoothness in Measure Spaces, with application to Sinkhorn and EMPierre-Cyril Aubin-Frankowski, Anna Korba, Flavien LégerNeurIPS 2022 · 61 citations
