Fault Lines: Benchmarking the Impact of Label Data Quality on ML Robustness and Fairness
David Jackson, Paul Groth, Hazar Harmouch
Abstract
Artificial intelligence systems depend critically on high-quality data, yet real-world datasets are often imperfect. Label noise, such as incorrect or biased labels, can lead to suboptimal model decisions. While label noise has garnered increasing attention, existing research primarily examines random noise, employs simpler models, or relies on limited evaluation criteria. To address this, we introduce Fault Lines, a comprehensive, model-agnostic benchmark comprising 15 datasets systematically corrupted with diverse types of label noise, paired with an evaluation framework. This resource supports the evaluation of data cleaning pipelines and guides the design of models that are robust, in both performance and fairness, to label noise. We benchmark the robustness to label noise of 22 state-of-the-art classification models, including gradient boosting, transformers, and fairness-oriented models. Our findings show that many models maintain strong performance under high random noise (e.g., up to 40% noise leads to only a modest reduction in Robust GBDT performance). However, these models are significantly less robust to even small amounts of biased noise (<10%), which can cause substantial performance drops (e.g., 7% noise reduces ResNet's AUC by 4.4% on average) or maintain apparent stability at the expense of severe fairness degradation (e.g., MLP's Predictive Parity difference increases by 700% under 30% biased noise in the ACS Unemployment dataset). We investigate how different model architectures handle the impact of biased noise. Notably, transformer-based models appear more robust than boosting models when handling biased noise, though this advantage depends on tuning and comes with higher variance. Finally, we identify key factors for ML practitioners to mitigate the effects of label noise, including model selection, dataset analysis, and preprocessing.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 62cd0035-2cd7-4cd2-a8e2-be01a0eceaaeBuilds on9
- Revisiting Deep Learning Models for Tabular DataYury Gorishniy, Ivan Rubachev, Valentin Khrulkov, Artem BabenkoNeurIPS 2021 · 1,847 citations
- DivideMix: Learning with Noisy Labels as Semi-supervised LearningJunnan Li, Richard Socher, Steven C. H. HoiICLR 2020 · 1,326 citations
- Out-of-Distribution Generalization via Risk Extrapolation (REx)David Krueger, Ethan Caballero, Jörn-Henrik Jacobsen, Amy Zhang et al.ICML 2021 · 1,163 citations
- Retiring Adult: New Datasets for Fair Machine LearningFrances Ding, Moritz Hardt, John Miller, Ludwig SchmidtNeurIPS 2021 · 671 citations
- Neural Oblivious Decision Ensembles for Deep Learning on Tabular DataSergei Popov, Stanislav Morozov, Artem BabenkoICLR 2020 · 407 citations
Related papers
- Stress-Testing ML Pipelines with Adversarial Data CorruptionJiongli Zhu, Geyang Xu, Felipe Lorenzi, Boris Glavic et al.VLDB 2025 · 2 citations
- AgentNoiseBench: Benchmarking Robustness of Tool-Using LLM Agents Under Noisy ConditionRuipeng Wang, Yuxin Chen, Yukai Wang, Chang Wu et al.ICML 2026 · 12 citations
- NoiseBench: Benchmarking the Impact of Real Label Noise on Named Entity RecognitionElena Merdjanovska, Ansar Aynetdinov, Alan AkbikEMNLP 2024 · 5 citations
- Dynamic Data Fault Localization for Deep Neural NetworksYining Yin, Yang Feng, Shihao Weng, Zixi Liu et al.FSE 2023 · 10 citations
- CleanML: A Study for Evaluating the Impact of Data Cleaning on ML Classification TasksPeng Li, Xi Rao, Jennifer Blase, Yue Zhang et al.ICDE 2021 · 127 citations
