On Saving Outliers for Better Clustering over Noisy Data
Shaoxu Song, Fei Gao, Ruihong Huang, Yihan Wang
Abstract
Clustering is often distracted by errors, frequently observed in almost all areas, ranging from online questionnaire to sensor reading in IoT. The dirty data values not only make themselves (the corresponding tuples) outlying, but also mislead the clustering of remaining tuples, e.g., mistakenly splitting a cluster into two or distorting the cluster center. The reason is that the traditional clustering methods either simply ignore the outliers such as DBSCAN or assign them to the closest clusters anyway, e.g., in K-Means. In this paper, we propose to save the outliers for better clustering. The idea is to adjust the erroneous values (often minimally) of the outlier in order to make it appear normally. That is, the tuples after adjusting values are no longer outlying, and thus will be clustered without distracting others. The outlier saving by value adjustment is designed to work with any clustering methods (e.g., DBSCAN or K-Means). Our technical contributions include: (1) showing NPhardness of the outlier saving problem for clustering, (2) deriving lower and upper bounds of the optimal solutions, and (3) devising approximation algorithm with performance guarantees referring to the aforesaid bounds. Experiments on datasets with real-world outliers demonstrate the higher accuracy of our proposal, compared to the state-of-the-art approaches. Remarkably, we show that the adjusted data with outlier saving indeed improve significantly clustering, as well as other applications such as classification and record matching.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- ShadowAQP: Efficient Approximate Group-by and Join Query via Attribute-oriented Sample Size Allocation and Data GenerationRong Gu, Han Li, Haipeng Dai, Wenjie Huang et al.VLDB 2023 · 9 citations
- Win-Win: On Simultaneous Clustering and Imputing over Incomplete DataYu Sun, Jingyu Zhu, Xiao Xu, Xian Xu et al.VLDB 2024 · 2 citations
- Efficient Structural Clustering Over HypergraphsDong Pan, Xu Zhou, Lingwei Li, Quanqing Xu et al.ICDE 2025
Related papers
- On Repairing Timestamps for Regular Interval Time SeriesChenguang Fang, Shaoxu Song, Yinan MeiVLDB 2022 · 18 citations
- The Best of Both Worlds: On Repairing Timestamps and Attribute Values for Multivariate Time SeriesJingyu Zhu, Weiwei Deng, Yu Sun, Shaoxu Song et al.SIGMOD 2025
- Cleaning Time Series under Seasonal and Trend ConstraintsZijie Chen, Aoqian Zhang, Shaoxu SongSIGMOD 2026
- Near-Linear Time Approximation Algorithms for k-means with OutliersJunyu Huang, Qilong Feng, Ziyun Huang, Jinhui Xu et al.ICML 2024 · 5 citations
- Fast Algorithms for Distributed k-Clustering with OutliersJunyu Huang, Qilong Feng, Ziyun Huang, Jinhui Xu et al.ICML 2023 · 7 citations
