SAGA: A Scalable Framework for Optimizing Data Cleaning Pipelines for Machine Learning Applications
Shafaq Siddiqi, Roman Kern, Matthias Boehm
Abstract
In the exploratory data science lifecycle, data scientists often spent the majority of their time finding, integrating, validating and cleaning relevant datasets. Despite recent work on data validation, and numerous error detection and correction algorithms, in practice, data cleaning for ML remains largely a manual, unpleasant, and labor-intensive trial and error process, especially in large-scale, distributed computation. The target ML application---such as classification or regression models---can be used as a signal of valuable feedback though, for selecting effective data cleaning strategies. In this paper, we introduce SAGA, a framework for automatically generating the top-K most effective data cleaning pipelines. SAGA adopts ideas from Auto-ML, feature selection, and hyper-parameter tuning. Our framework is extensible for user-provided constraints, new data cleaning primitives, and ML applications; automatically generates hybrid runtime plans of local and distributed operations; and performs pruning by interesting properties (e.g., monotonicity). Instead of full automation---which is rather unrealistic---SAGA simplifies the mechanical aspects of data cleaning. Our experiments show that SAGA yields robust accuracy improvements over state-of-the-art, and good scalability regarding increasing data sizes and number of evaluated pipelines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e24d13e9-6c2a-4307-ab6f-d8f4da67cc67Cited by top-tier papers6
- Malleus: Straggler-Resilient Hybrid Parallel Training of Large-scale Models via Malleable Data and Model ParallelizationHaoyang Li, Fangcheng Fu, Hao Ge, Sheng Lin et al.SIGMOD 2025 · 6 citations
- CatDB: Data-catalog-guided, LLM-based Generation of Data-centric ML PipelinesSaeed Fathollahzadeh, Essam Mansour, Matthias BoehmVLDB 2025 · 5 citations
- DeMix: Debugging Training Data with Mixed Data Error Types by Investigating Influence VectorsJiale Deng, Yanyan Shen, Xiaogang Shi, Junjun ChaiKDD 2026
- Morphing-based Compression for Data-centric ML PipelinesSebastian Baunsgaard, Matthias BoehmVLDB 2026
- Understanding the Impact of Data Noise in Federated Learning: [Experiments & Analysis]Jinming Hu, Jiahao Gu, Kenta Ploch, Hao Wang et al.SIGMOD 2026
Builds on12
- CleanML: A Study for Evaluating the Impact of Data Cleaning on ML Classification TasksPeng Li, Xi Rao, Jennifer Blase, Yue Zhang et al.ICDE 2021 · 127 citations
- Nearest Neighbor Classifiers over Incomplete Information: From Certain Answers to Certain PredictionsBojan Karlas, Peng Li, Renzhi Wu, Nezihe Merve Gürel et al.VLDB 2021 · 69 citations
- Cerebro: A Data System for Optimized Deep Learning Model SelectionSupun Nakandala, Yuhao Zhang, Arun KumarVLDB 2020 · 61 citations
- SliceLine: Fast, Linear-Algebra-based Slice Finding for ML Model DebuggingSvetlana Sagadeeva, Matthias BoehmSIGMOD 2021 · 45 citations
- Automated Feature Engineering for Algorithmic FairnessRicardo Salazar, Felix Neutatz, Ziawasch AbedjanVLDB 2021 · 42 citations
Related papers
- LIMA: Fine-grained Lineage Tracing and Reuse in Machine Learning SystemsArnab Phani, Benjamin Rath, Matthias BoehmSIGMOD 2021 · 30 citations
- DMCO: Budget-Aware Co-Optimization of Data Cleaning and AutoMLXiaoou Ding, Zekai Qian, Siying Chen, Hongbin Hu et al.ICML 2026
- AutoDS: Towards Human-Centered Automation of Data ScienceDakuo Wang, Josh Andres, Justin D. Weisz, Erick Oduor et al.CHI 2021 · 77 citations
- SPIO: Ensemble and Selective Strategies via LLM-Based Multi-Agent Planning in Automated Data ScienceWonduk Seo, Juhyeon Lee, Yanjun Shao, Qingshan Zhou et al.ACL 2026 · 8 citations
- SAPIENTML: Synthesizing Machine Learning Pipelines by Learning from Human-Written SolutionsRipon K. Saha, Akira Ura, Sonal Mahajan, Chenguang Zhu et al.ICSE 2022 · 11 citations
