On Saving Outliers for Better Clustering over Noisy Data
Shaoxu Song, Fei Gao, Ruihong Huang, Yihan Wang
摘要
Clustering is often distracted by errors, frequently observed in almost all areas, ranging from online questionnaire to sensor reading in IoT. The dirty data values not only make themselves (the corresponding tuples) outlying, but also mislead the clustering of remaining tuples, e.g., mistakenly splitting a cluster into two or distorting the cluster center. The reason is that the traditional clustering methods either simply ignore the outliers such as DBSCAN or assign them to the closest clusters anyway, e.g., in K-Means. In this paper, we propose to save the outliers for better clustering. The idea is to adjust the erroneous values (often minimally) of the outlier in order to make it appear normally. That is, the tuples after adjusting values are no longer outlying, and thus will be clustered without distracting others. The outlier saving by value adjustment is designed to work with any clustering methods (e.g., DBSCAN or K-Means). Our technical contributions include: (1) showing NPhardness of the outlier saving problem for clustering, (2) deriving lower and upper bounds of the optimal solutions, and (3) devising approximation algorithm with performance guarantees referring to the aforesaid bounds. Experiments on datasets with real-world outliers demonstrate the higher accuracy of our proposal, compared to the state-of-the-art approaches. Remarkably, we show that the adjusted data with outlier saving indeed improve significantly clustering, as well as other applications such as classification and record matching.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- ShadowAQP: Efficient Approximate Group-by and Join Query via Attribute-oriented Sample Size Allocation and Data GenerationRong Gu, Han Li, Haipeng Dai, Wenjie Huang 等VLDB 2023 · 被引用 9 次
- Win-Win: On Simultaneous Clustering and Imputing over Incomplete DataYu Sun, Jingyu Zhu, Xiao Xu, Xian Xu 等VLDB 2024 · 被引用 2 次
- Efficient Structural Clustering Over HypergraphsDong Pan, Xu Zhou, Lingwei Li, Quanqing Xu 等ICDE 2025
相关 Paper
- On Repairing Timestamps for Regular Interval Time SeriesChenguang Fang, Shaoxu Song, Yinan MeiVLDB 2022 · 被引用 18 次
- The Best of Both Worlds: On Repairing Timestamps and Attribute Values for Multivariate Time SeriesJingyu Zhu, Weiwei Deng, Yu Sun, Shaoxu Song 等SIGMOD 2025
- Cleaning Time Series under Seasonal and Trend ConstraintsZijie Chen, Aoqian Zhang, Shaoxu SongSIGMOD 2026
- Near-Linear Time Approximation Algorithms for k-means with OutliersJunyu Huang, Qilong Feng, Ziyun Huang, Jinhui Xu 等ICML 2024 · 被引用 5 次
- Fast Algorithms for Distributed k-Clustering with OutliersJunyu Huang, Qilong Feng, Ziyun Huang, Jinhui Xu 等ICML 2023 · 被引用 7 次
