Automatic Feasibility Study via Data Quality Analysis for ML: A Case-Study on Label Noise
Cédric Renggli, Luka Rimanic, Luka Kolar, Wentao Wu, Ce Zhang
Abstract
In our experience of working with domain experts who are using today’s AutoML systems, a common problem we encountered is what we call "unrealistic expectations" – when users are facing a very challenging task with a noisy data acquisition process, while being expected to achieve startlingly high accuracy with machine learning (ML). Many of these are predestined to fail from the beginning. In traditional software engineering, this problem is addressed via a feasibility study, an indispensable step before developing any software system. In this paper, we present Snoopy, with the goal of supporting data scientists and machine learning engineers performing a systematic and theoretically founded feasibility study before building ML applications. We approach this problem by estimating the irreducible error of the underlying task, also known as the Bayes error rate (BER), which stems from data quality issues in datasets used to train or evaluate ML models. We design a practical Bayes error estimator that is compared against baseline feasibility study candidates on 6 datasets (with additional real and synthetic noise of different levels) in computer vision and natural language processing. Furthermore, by including our systematic feasibility study with additional signals into the iterative label cleaning process, we demonstrate in end-to-end experiments how users are able to save substantial labeling time and monetary efforts.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- T-Assess: An Efficient Data Quality Assessment System Tailored for Trajectory DataJunhao Zhu, Tao Wang, Danlei Hu, Ziquan Fang et al.VLDB 2025 · 2 citations
- Measuring Database Unfairness via Dependency Quantification Under Differential PrivacyMariia Vologdin, Yuchao Tao, Amir GiladVLDB 2026
- Understanding the Impact of Data Noise in Federated Learning: [Experiments & Analysis]Jinming Hu, Jiahao Gu, Kenta Ploch, Hao Wang et al.SIGMOD 2026
Builds on11
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
- Learning with Noisy Labels Revisited: A Study Using Real-World Human AnnotationsJiaheng Wei, Zhaowei Zhu, Hao Cheng, Tongliang Liu et al.ICLR 2022 · 338 citations
- A Constructive Prediction of the Generalization Error Across ScalesJonathan S. Rosenfeld, Amir Rosenfeld, Yonatan Belinkov, Nir ShavitICLR 2020 · 265 citations
- CleanML: A Study for Evaluating the Impact of Data Cleaning on ML Classification TasksPeng Li, Xi Rao, Jennifer Blase, Yue Zhang et al.ICDE 2021 · 127 citations
Related papers
- Demystifying the Optimal Performance of Multi-Class ClassificationMinoh Jeong, Martina Cardone, Alex DytsoNeurIPS 2023 · 17 citations
- BClean: A Bayesian Data Cleaning SystemJianbin Qin, Sifan Huang, Yaoshu Wang, Jing Zhu et al.ICDE 2024 · 4 citations
- Is the Performance of My Deep Network Too Good to Be True? A Direct Approach to Estimating the Bayes Error in Binary ClassificationTakashi Ishida, Ikko Yamane, Nontawat Charoenphakdee, Gang Niu et al.ICLR 2023 · 3 citations
- Physics-Driven ML-Based Modelling for Correcting Inverse EstimationRuiyuan Kang, Tingting Mu, Panagiotis Liatsis, Dimitrios C. KyritsisNeurIPS 2023 · 2 citations
- Detecting Label Errors by Using Pre-Trained Language ModelsDerek Chong, Jenny Hong, Christopher D. ManningEMNLP 2022 · 8 citations
