CleanML: A Study for Evaluating the Impact of Data Cleaning on ML Classification Tasks
Peng Li, Xi Rao, Jennifer Blase, Yue Zhang, Xu Chu, Ce Zhang
摘要
Data quality affects machine learning (ML) model performances, and data scientists spend considerable amount of time on data cleaning before model training. However, to date, there does not exist a rigorous study on how exactly cleaning affects ML - ML community usually focuses on developing ML algorithms that are robust to some particular noise types of certain distributions, while database (DB) community has been mostly studying the problem of data cleaning alone without considering how data is consumed by downstream ML analytics.We propose a CleanML study that systematically investigates the impact of data cleaning on ML classification tasks. The open-source and extensible CleanML study currently includes 14 real-world datasets with real errors, five common error types, seven different ML models, and multiple cleaning algorithms for each error type (including both commonly used algorithms in practice as well as state-of-the-art solutions in academic literature). We control the randomness in ML experiments using statistical hypothesis testing, and we also control false discovery rate in our experiments using the Benjamini-Yekutieli (BY) procedure. We analyze the results in a systematic way to derive many interesting and nontrivial observations. We also put forward multiple research directions for researchers.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper26
- GoodCore: Data-effective and Data-efficient Machine Learning through Coreset Selection over Incomplete DataChengliang Chai, Jiabin Liu, Nan Tang, Ju Fan 等SIGMOD 2023 · 被引用 37 次
- Entity Resolution On-DemandGiovanni Simonini, Luca Zecchini, Sonia Bergamaschi, Felix NaumannVLDB 2022 · 被引用 33 次
- Sudowoodo: Contrastive Self-supervised Learning for Multi-purpose Data Integration and PreparationRunhui Wang, Yuliang Li, Jin WangICDE 2023 · 被引用 32 次
- Automatic Data Repair: Are We Ready to Deploy?Wei Ni, Xiaoye Miao, Xiangyu Zhao, Yangyang Wu 等VLDB 2024 · 被引用 26 次
- DiffPrep: Differentiable Data Preprocessing Pipeline Search for Learning over Tabular DataPeng Li, Zhiyi Chen, Xu Chu, Kexin RongSIGMOD 2023 · 被引用 24 次
它引用的顶会 Paper2
相关 Paper
- DeMix: Debugging Training Data with Mixed Data Error Types by Investigating Influence VectorsJiale Deng, Yanyan Shen, Xiaogang Shi, Junjun ChaiKDD 2026
- Are we building on the rock? on the importance of data preprocessing for code summarizationLin Shi, Fangwen Mu, Xiao Chen, Song Wang 等FSE 2022 · 被引用 65 次
- SAGA: A Scalable Framework for Optimizing Data Cleaning Pipelines for Machine Learning ApplicationsShafaq Siddiqi, Roman Kern, Matthias BoehmSIGMOD 2024 · 被引用 24 次
- Automatic Feasibility Study via Data Quality Analysis for ML: A Case-Study on Label NoiseCédric Renggli, Luka Rimanic, Luka Kolar, Wentao Wu 等ICDE 2023 · 被引用 8 次
- How do Categorical Duplicates Affect ML? A New Benchmark and Empirical AnalysesVraj Shah, Thomas J. Parashos, Arun KumarVLDB 2024 · 被引用 8 次
