Stress-Testing ML Pipelines with Adversarial Data Corruption
Jiongli Zhu, Geyang Xu, Felipe Lorenzi, Boris Glavic, Babak Salimi
Abstract
Structured data-quality issues—such as missing values correlated with demographics, culturally biased labels, or systemic selection biases—routinely degrade the reliability of machine-learning pipelines. Regulators now increasingly demand evidence that high-stakes systems can withstand these realistic, interdependent errors, yet current robustness evaluations typically use random or overly simplistic corruptions, leaving worst-case scenarios unexplored. We introduce Savage, a causally inspired framework that (i) formally models realistic data-quality issues through dependency graphs and flexible corruption templates, and (ii) systematically discovers corruption patterns that maximally degrade a target performance metric. Savage employs a bi-level optimization approach to efficiently identify vulnerable data subpopulations and fine-tune corruption severity, treating the full ML pipeline, including preprocessing and potentially non-differentiable models, as a black box. Extensive experiments across multiple datasets and ML tasks (data cleaning, fairness-aware learning, uncertainty quantification) demonstrate that even a small fraction (around 5%) of structured corruptions identified by Savage severely impacts model performance, far exceeding random or manually crafted errors, and invalidating core assumptions of existing techniques. Thus, Savage provides a practical tool for rigorous pipeline stress-testing, a benchmark for evaluating robustness methods, and actionable guidance for designing more resilient data workflows.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2cdf8736-5702-4aba-98ae-d6cf2d376016Builds on12
- Just How Toxic is Data Poisoning? A Unified Benchmark for Backdoor and Data Poisoning AttacksAvi Schwarzschild, Micah Goldblum, Arjun Gupta, John P. Dickerson et al.ICML 2021 · 207 citations
- CleanML: A Study for Evaluating the Impact of Data Cleaning on ML Classification TasksPeng Li, Xi Rao, Jennifer Blase, Yue Zhang et al.ICDE 2021 · 127 citations
- Sample Selection for Fair and Robust TrainingYuji Roh, Kangwook Lee, Steven Whang, Changho SuhNeurIPS 2021 · 76 citations
- Hidden Poison: Machine Unlearning Enables Camouflaged Poisoning AttacksJimmy Z. Di, Jack Douglas, Jayadev Acharya, Gautam Kamath et al.NeurIPS 2023 · 70 citations
- Interpretable Data-Based Explanations for Fairness DebuggingRomila Pradhan, Jiongli Zhu, Boris Glavic, Babak SalimiSIGMOD 2022 · 53 citations
Related papers
- Stress-Testing Causal Claims via Cardinality RepairsYarden Gabbay, Haoquan Guan, Shaull Almagor, El Kindi Rezig et al.SIGMOD 2026 · 1 citation
- Deception by Omission: Using Adversarial Missingness to Poison Causal Structure LearningDeniz Koyuncu, Alex Gittens, Bülent Yener, Moti YungKDD 2023 · 1 citation
- SAGA: A Scalable Framework for Optimizing Data Cleaning Pipelines for Machine Learning ApplicationsShafaq Siddiqi, Roman Kern, Matthias BoehmSIGMOD 2024 · 24 citations
- (Be Cautious!) Bio-Foundation Models Are Not Yet Robust to Biologically Plausible Perturbations and ML TransformationsJinhao Duan, Ruichen Zhang, Gengwei Zhang, Huaizhi Qu et al.ICML 2026
- Mind the Graph When Balancing Data for Fairness or RobustnessJessica Schrouff, Alexis Bellot, Amal Rannen-Triki, Alan Malek et al.NeurIPS 2024 · 10 citations
