SAGA: A Scalable Framework for Optimizing Data Cleaning Pipelines for Machine Learning Applications
Shafaq Siddiqi, Roman Kern, Matthias Boehm
摘要
In the exploratory data science lifecycle, data scientists often spent the majority of their time finding, integrating, validating and cleaning relevant datasets. Despite recent work on data validation, and numerous error detection and correction algorithms, in practice, data cleaning for ML remains largely a manual, unpleasant, and labor-intensive trial and error process, especially in large-scale, distributed computation. The target ML application---such as classification or regression models---can be used as a signal of valuable feedback though, for selecting effective data cleaning strategies. In this paper, we introduce SAGA, a framework for automatically generating the top-K most effective data cleaning pipelines. SAGA adopts ideas from Auto-ML, feature selection, and hyper-parameter tuning. Our framework is extensible for user-provided constraints, new data cleaning primitives, and ML applications; automatically generates hybrid runtime plans of local and distributed operations; and performs pruning by interesting properties (e.g., monotonicity). Instead of full automation---which is rather unrealistic---SAGA simplifies the mechanical aspects of data cleaning. Our experiments show that SAGA yields robust accuracy improvements over state-of-the-art, and good scalability regarding increasing data sizes and number of evaluated pipelines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Malleus: Straggler-Resilient Hybrid Parallel Training of Large-scale Models via Malleable Data and Model ParallelizationHaoyang Li, Fangcheng Fu, Hao Ge, Sheng Lin 等SIGMOD 2025 · 被引用 6 次
- CatDB: Data-catalog-guided, LLM-based Generation of Data-centric ML PipelinesSaeed Fathollahzadeh, Essam Mansour, Matthias BoehmVLDB 2025 · 被引用 5 次
- DeMix: Debugging Training Data with Mixed Data Error Types by Investigating Influence VectorsJiale Deng, Yanyan Shen, Xiaogang Shi, Junjun ChaiKDD 2026
- Morphing-based Compression for Data-centric ML PipelinesSebastian Baunsgaard, Matthias BoehmVLDB 2026
- Understanding the Impact of Data Noise in Federated Learning: [Experiments & Analysis]Jinming Hu, Jiahao Gu, Kenta Ploch, Hao Wang 等SIGMOD 2026
它引用的顶会 Paper12
- CleanML: A Study for Evaluating the Impact of Data Cleaning on ML Classification TasksPeng Li, Xi Rao, Jennifer Blase, Yue Zhang 等ICDE 2021 · 被引用 127 次
- Nearest Neighbor Classifiers over Incomplete Information: From Certain Answers to Certain PredictionsBojan Karlas, Peng Li, Renzhi Wu, Nezihe Merve Gürel 等VLDB 2021 · 被引用 69 次
- Cerebro: A Data System for Optimized Deep Learning Model SelectionSupun Nakandala, Yuhao Zhang, Arun KumarVLDB 2020 · 被引用 61 次
- SliceLine: Fast, Linear-Algebra-based Slice Finding for ML Model DebuggingSvetlana Sagadeeva, Matthias BoehmSIGMOD 2021 · 被引用 45 次
- Automated Feature Engineering for Algorithmic FairnessRicardo Salazar, Felix Neutatz, Ziawasch AbedjanVLDB 2021 · 被引用 42 次
相关 Paper
- LIMA: Fine-grained Lineage Tracing and Reuse in Machine Learning SystemsArnab Phani, Benjamin Rath, Matthias BoehmSIGMOD 2021 · 被引用 30 次
- DMCO: Budget-Aware Co-Optimization of Data Cleaning and AutoMLXiaoou Ding, Zekai Qian, Siying Chen, Hongbin Hu 等ICML 2026
- AutoDS: Towards Human-Centered Automation of Data ScienceDakuo Wang, Josh Andres, Justin D. Weisz, Erick Oduor 等CHI 2021 · 被引用 77 次
- SPIO: Ensemble and Selective Strategies via LLM-Based Multi-Agent Planning in Automated Data ScienceWonduk Seo, Juhyeon Lee, Yanjun Shao, Qingshan Zhou 等ACL 2026 · 被引用 8 次
- SAPIENTML: Synthesizing Machine Learning Pipelines by Learning from Human-Written SolutionsRipon K. Saha, Akira Ura, Sonal Mahajan, Chenguang Zhu 等ICSE 2022 · 被引用 11 次
