Likelihood Adjusted Semidefinite Programs for Clustering Heterogeneous Data
Yubo Zhuang, Xiaohui Chen, Yun Yang
Abstract
Clustering is a widely deployed unsupervised learning tool. Model-based clustering is a flexible framework to tackle data heterogeneity when the clusters have different shapes. Likelihood-based inference for mixture distributions often involves non-convex and high-dimensional objective functions, imposing difficult computational and statistical challenges. The classic expectation-maximization (EM) algorithm is a computationally thrifty iterative method that maximizes a surrogate function minorizing the log-likelihood of observed data in each iteration, which however suffers from bad local maxima even in the special case of the standard Gaussian mixture model with common isotropic covariance matrices. On the other hand, recent studies reveal that the unique global solution of a semidefinite programming (SDP) relaxed -means achieves the information-theoretically sharp threshold for perfectly recovering the cluster labels under the standard Gaussian mixture model. In this paper, we extend the SDP approach to a general setting by integrating cluster labels as model parameters and propose an iterative likelihood adjusted SDP (iLA-SDP) method that directly maximizes the exact observed likelihood in the presence of data heterogeneity. By lifting the cluster assignment to group-specific membership matrices, iLA-SDP avoids centroids estimation -- a key feature that allows exact recovery under well-separateness of centroids without being trapped by their adversarial configurations. Thus iLA-SDP is less sensitive than EM to initialization and more stable on high-dimensional data. Our numeric experiments demonstrate that iLA-SDP can achieve lower mis-clustering errors over several widely used clustering methods including -means, SDP and EM algorithms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1a7863d1-553d-43c8-bf56-c2b7a2b065aeCited by top-tier papers1
Ask how each one uses itBuilds on2
Related papers
- Mixtures Closest To A Given Measure: A Semidefinite Programming ApproachSrećko Ðurašinović, Jean B Lasserre, Victor MagronICML 2026
- Achieving Optimal Clustering in Gaussian Mixture Models with Anisotropic Covariance StructuresXin Chen, Anderson Ye ZhangNeurIPS 2024 · 15 citations
- A Federated Generalized Expectation-Maximization Algorithm for Mixture Models with an Unknown Number of ComponentsMichael Ibrahim, Nagi Gebraeel, Weijun XieICLR 2026 · 1 citation
- Big Learning Expectation MaximizationYulai Cong, Sijia LiAAAI 2024 · 5 citations
- Heterogeneous Sufficient Dimension Reduction and Subspace ClusteringLei Yan, Xin Zhang, Qing MaiICML 2025
