Toward Global Convergence of Gradient EM for Over-Paramterized Gaussian Mixture Models
Weihang Xu, Maryam Fazel, Simon S. Du
Abstract
We study the gradient Expectation-Maximization (EM) algorithm for Gaussian Mixture Models (GMM) in the over-parameterized setting, where a general GMM with components learns from data that are generated by a single ground truth Gaussian distribution. While results for the special case of 2-Gaussian mixtures are well-known, a general global convergence analysis for arbitrary remains unresolved and faces several new technical barriers since the convergence becomes sub-linear and non-monotonic. To address these challenges, we construct a novel likelihood-based convergence analysis framework and rigorously prove that gradient EM converges globally with a sublinear rate . This is the first global convergence result for Gaussian mixtures with more than components. The sublinear convergence rate is due to the algorithmic nature of learning over-parameterized GMM with gradient EM. We also identify a new emerging technical challenge for learning general over-parameterized GMM: the existence of bad local regions that can trap gradient EM for an exponential number of steps.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c5f35c26-c1a9-48de-ad23-669b3612ac46Cited by top-tier papers1
Ask how each one uses itBuilds on2
- Preconditioned Gradient Descent for Over-Parameterized Nonconvex Matrix FactorizationJialun Zhang, Salar Fattahi, Richard Y. ZhangNeurIPS 2021 · 47 citations
- How Over-Parameterization Slows Down Gradient Descent in Matrix Sensing: The Curses of Symmetry and InitializationNuoya Xiong, Lijun Ding, Simon Shaolei DuICLR 2024 · 22 citations
Related papers
- Mean Estimation of Truncated Mixtures of Two Gaussians: A Gradient Based ApproachSai Ganesh Nagarajan, Gerasimos Palaiopanos, Ioannis Panageas, Tushar Vaidya et al.AAAI 2023
- Unveiling the Cycloid Trajectory of EM Iterations in Mixed Linear RegressionZhankun Luo, Abolfazl HashemiICML 2024 · 1 citation
- A Federated Generalized Expectation-Maximization Algorithm for Mixture Models with an Unknown Number of ComponentsMichael Ibrahim, Nagi Gebraeel, Weijun XieICLR 2026 · 1 citation
- Learning Mixtures of Gaussians Using the DDPM ObjectiveKulin Shah, Sitan Chen, Adam R. KlivansNeurIPS 2023 · 69 citations
- Learning Mixtures of Experts with EM: A Mirror Descent PerspectiveQuentin Fruytier, Aryan Mokhtari, Sujay SanghaviICML 2025
