Win-Win: On Simultaneous Clustering and Imputing over Incomplete Data
Yu Sun, Jingyu Zhu, Xiao Xu, Xian Xu, Yuyao Sun, Shaoxu Song, Xiang Li, Xiaojie Yuan
Abstract
Although clustering methods have shown promising performance in various applications, they cannot effectively handle incomplete data. Existing studies often impute missing values first before clustering analysis and conduct these two processes separately. However, inaccurate imputation does not necessarily contribute positively to the subsequent clustering. Intuitively, accurate imputation and clustering can serve and benefit from each other, where clustering-based imputation methods typically utilize cluster signals to impute incomplete data and accurate fillings are expected to bring more valuable data for clustering. Therefore, in this manuscript, rather than considering two tasks independently or conducting them respectively, we study simultaneous clustering and imputing over incomplete data. The immediate benefit is that such a strategy improves both clustering and imputation performance simultaneously, to get a win-win result. Our major technical highlights include (1) the problem formalization and NP-hardness analysis on computing simultaneous clustering and imputing results, (2) exact solutions by transforming the problem as the integer linear programming (ILP) formulation, and (3) efficient approximation algorithms based on the linear programming (LP) relaxation and local neighbors (LN) solution, with approximation guarantees. Experiments on various real-world datasets demonstrate the superiority of our work in clustering and imputing incomplete data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on4
- CSDI: Conditional Score-based Diffusion Models for Probabilistic Time Series ImputationYusuke Tashiro, Jiaming Song, Yang Song, Stefano ErmonNeurIPS 2021 · 1,245 citations
- Nearest Neighbor Classifiers over Incomplete Information: From Certain Answers to Certain PredictionsBojan Karlas, Peng Li, Renzhi Wu, Nezihe Merve Gürel et al.VLDB 2021 · 69 citations
- Imputing Various Incomplete Attributes via Distance Likelihood MaximizationShaoxu Song, Yu SunKDD 2020 · 15 citations
- On Saving Outliers for Better Clustering over Noisy DataShaoxu Song, Fei Gao, Ruihong Huang, Yihan WangSIGMOD 2021 · 6 citations
Related papers
- Attribute-Missing Graph Clustering NetworkWenxuan Tu, Renxiang Guan, Sihang Zhou, Chuan Ma et al.AAAI 2024 · 51 citations
- Deep Safe Incomplete Multi-view Clustering: Theorem and AlgorithmHuayi Tang, Yong LiuICML 2022 · 118 citations
- Learning Representations for Incomplete Time Series ClusteringQianli Ma, Chuxin Chen, Sen Li, Garrison W. CottrellAAAI 2021 · 34 citations
- Incomplete Multi-view Deep Clustering with Data Imputation and AlignmentJiyuan Liu, Xinwang Liu, Xinhang Wan, Ke Liang et al.NeurIPS 2025 · 1 citation
- Missing Value Imputation for Mixed Data via Gaussian CopulaYuxuan Zhao, Madeleine UdellKDD 2020 · 53 citations
