Outliers: The Good, the Bad and the Ugly
Shenglin Chen, Wenfei Fan, Ruochun Jin
摘要
This paper studies the impact of outliers in relational dataset D on the accuracy of ML classifiers M , when D is used to train and evaluate M . Outliers are data points that statistically deviate from the distribution of the majority. We distinguish good outliers, i.e. novel data, from bad ones, i.e. those introduced by errors. Moreover, we separate ugly ones in influential features from the other bad ones. We find that only the ugly ones have negative impact on M , while the good (resp. the other bad) ones have positive (resp. neglectable) impact. To mitigate the negative impact, we propose a class of rules, denoted by OMRs, to identify ugly outliers by embedding ML outlier detectors and statistical functions as predicates. We develop algorithms to (a) learn OMRs from real-life data, and (b) catch and fix ugly outliers using the learned OMRs, instead of removing tuples. Using real-life data, we empirically show that OMRs improve the accuracy of various classifiers by 7.2% on average, up to 34.8%.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Neural Relation Graph: A Unified Framework for Identifying Label Noise and Outlier DataJang-Hyun Kim, Sangdoo Yun, Hyun Oh SongNeurIPS 2023 · 被引用 32 次
- Data Enhancement for Binary Classification of Relational DataWenfei Fan, Xiaoyu Han, Weilong Ren, Zihuan XuSIGMOD 2025 · 被引用 1 次
- Conflict Resolution for Improving ML AccuracyWenfei Fan, Xiaoyu Han, Hufsa Khan, Weilong Ren 等ICDE 2026
- Data Glitches Discovery using Influence-based Model ExplanationsNikolaos Myrtakis, Ioannis Tsamardinos, Vassilis ChristophidesKDD 2025
- CleanML: A Study for Evaluating the Impact of Data Cleaning on ML Classification TasksPeng Li, Xi Rao, Jennifer Blase, Yue Zhang 等ICDE 2021 · 被引用 127 次
