Likelihood Adjusted Semidefinite Programs for Clustering Heterogeneous Data
Yubo Zhuang, Xiaohui Chen, Yun Yang
摘要
Clustering is a widely deployed unsupervised learning tool. Model-based clustering is a flexible framework to tackle data heterogeneity when the clusters have different shapes. Likelihood-based inference for mixture distributions often involves non-convex and high-dimensional objective functions, imposing difficult computational and statistical challenges. The classic expectation-maximization (EM) algorithm is a computationally thrifty iterative method that maximizes a surrogate function minorizing the log-likelihood of observed data in each iteration, which however suffers from bad local maxima even in the special case of the standard Gaussian mixture model with common isotropic covariance matrices. On the other hand, recent studies reveal that the unique global solution of a semidefinite programming (SDP) relaxed -means achieves the information-theoretically sharp threshold for perfectly recovering the cluster labels under the standard Gaussian mixture model. In this paper, we extend the SDP approach to a general setting by integrating cluster labels as model parameters and propose an iterative likelihood adjusted SDP (iLA-SDP) method that directly maximizes the exact observed likelihood in the presence of data heterogeneity. By lifting the cluster assignment to group-specific membership matrices, iLA-SDP avoids centroids estimation -- a key feature that allows exact recovery under well-separateness of centroids without being trapped by their adversarial configurations. Thus iLA-SDP is less sensitive than EM to initialization and more stable on high-dimensional data. Our numeric experiments demonstrate that iLA-SDP can achieve lower mis-clustering errors over several widely used clustering methods including -means, SDP and EM algorithms.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper2
相关 Paper
- Mixtures Closest To A Given Measure: A Semidefinite Programming ApproachSrećko Ðurašinović, Jean B Lasserre, Victor MagronICML 2026
- Achieving Optimal Clustering in Gaussian Mixture Models with Anisotropic Covariance StructuresXin Chen, Anderson Ye ZhangNeurIPS 2024 · 被引用 15 次
- A Federated Generalized Expectation-Maximization Algorithm for Mixture Models with an Unknown Number of ComponentsMichael Ibrahim, Nagi Gebraeel, Weijun XieICLR 2026 · 被引用 1 次
- Big Learning Expectation MaximizationYulai Cong, Sijia LiAAAI 2024 · 被引用 5 次
- Heterogeneous Sufficient Dimension Reduction and Subspace ClusteringLei Yan, Xin Zhang, Qing MaiICML 2025
