Win-Win: On Simultaneous Clustering and Imputing over Incomplete Data
Yu Sun, Jingyu Zhu, Xiao Xu, Xian Xu, Yuyao Sun, Shaoxu Song, Xiang Li, Xiaojie Yuan
摘要
Although clustering methods have shown promising performance in various applications, they cannot effectively handle incomplete data. Existing studies often impute missing values first before clustering analysis and conduct these two processes separately. However, inaccurate imputation does not necessarily contribute positively to the subsequent clustering. Intuitively, accurate imputation and clustering can serve and benefit from each other, where clustering-based imputation methods typically utilize cluster signals to impute incomplete data and accurate fillings are expected to bring more valuable data for clustering. Therefore, in this manuscript, rather than considering two tasks independently or conducting them respectively, we study simultaneous clustering and imputing over incomplete data. The immediate benefit is that such a strategy improves both clustering and imputation performance simultaneously, to get a win-win result. Our major technical highlights include (1) the problem formalization and NP-hardness analysis on computing simultaneous clustering and imputing results, (2) exact solutions by transforming the problem as the integer linear programming (ILP) formulation, and (3) efficient approximation algorithms based on the linear programming (LP) relaxation and local neighbors (LN) solution, with approximation guarantees. Experiments on various real-world datasets demonstrate the superiority of our work in clustering and imputing incomplete data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- CSDI: Conditional Score-based Diffusion Models for Probabilistic Time Series ImputationYusuke Tashiro, Jiaming Song, Yang Song, Stefano ErmonNeurIPS 2021 · 被引用 1,245 次
- Nearest Neighbor Classifiers over Incomplete Information: From Certain Answers to Certain PredictionsBojan Karlas, Peng Li, Renzhi Wu, Nezihe Merve Gürel 等VLDB 2021 · 被引用 69 次
- Imputing Various Incomplete Attributes via Distance Likelihood MaximizationShaoxu Song, Yu SunKDD 2020 · 被引用 15 次
- On Saving Outliers for Better Clustering over Noisy DataShaoxu Song, Fei Gao, Ruihong Huang, Yihan WangSIGMOD 2021 · 被引用 6 次
相关 Paper
- Attribute-Missing Graph Clustering NetworkWenxuan Tu, Renxiang Guan, Sihang Zhou, Chuan Ma 等AAAI 2024 · 被引用 51 次
- Deep Safe Incomplete Multi-view Clustering: Theorem and AlgorithmHuayi Tang, Yong LiuICML 2022 · 被引用 118 次
- Learning Representations for Incomplete Time Series ClusteringQianli Ma, Chuxin Chen, Sen Li, Garrison W. CottrellAAAI 2021 · 被引用 34 次
- Incomplete Multi-view Deep Clustering with Data Imputation and AlignmentJiyuan Liu, Xinwang Liu, Xinhang Wan, Ke Liang 等NeurIPS 2025 · 被引用 1 次
- Missing Value Imputation for Mixed Data via Gaussian CopulaYuxuan Zhao, Madeleine UdellKDD 2020 · 被引用 53 次
