DMCO: Budget-Aware Co-Optimization of Data Cleaning and AutoML
Xiaoou Ding, Zekai Qian, Siying Chen, Hongbin Hu, Chen Wang, Hongzhi Wang, Jianmin Wang
摘要
Data cleaning and automated machine learning (AutoML) are both crucial for reliable learning systems, yet are commonly treated as independent or sequential stages. This separation ignores their strong interaction and leads to inefficient use of limited computational budgets. We propose DMCO, a unified framework that jointly optimizes data cleaning and model construction under a fixed resource budget. DMCO reformulates the traditional two-stage pipeline into a time-sliced process, where data cleaning and AutoML are interleaved and adaptively scheduled. We introduce a gradient-based data cleaning sampling strategy with theoretical guarantees for minimizing gradient estimation variance, and integrates it with loss-driven sampling and progressive AutoML fitting to continuously leverage intermediate data quality improvements. Experiments on six realworld datasets show that DMCO consistently outperforms standalone data cleaning and AutoML baselines on both classification and regression tasks, as measured by F1 score and MSE. Under limited budgets, DMCO achieves up to 82.19% of the performance of full data cleaning with exhaustive AutoML, while remaining robust across different AutoML frameworks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- Horizon: Scalable Dependency-driven Data CleaningEl Kindi Rezig, Mourad Ouzzani, Walid G. Aref, Ahmed K. Elmagarmid 等VLDB 2021 · 被引用 95 次
- Nearest Neighbor Classifiers over Incomplete Information: From Certain Answers to Certain PredictionsBojan Karlas, Peng Li, Renzhi Wu, Nezihe Merve Gürel 等VLDB 2021 · 被引用 69 次
- MTSClean: Efficient Constraint-based Cleaning for Multi-Dimensional Time Series DataXiaoou Ding, Yichen Song, Hongzhi Wang, Chen Wang 等VLDB 2024 · 被引用 9 次
- UniClean: A Scalable Data Cleaning Solution for Mixed Errors based on Unified Cleaners and Optimized Cleaning WorkflowXiaoou Ding, Zekai Qian, Hongzhi Wang, Siying Chen 等VLDB 2025 · 被引用 1 次
相关 Paper
- DiffPrep: Differentiable Data Preprocessing Pipeline Search for Learning over Tabular DataPeng Li, Zhiyi Chen, Xu Chu, Kexin RongSIGMOD 2023 · 被引用 24 次
- SAGA: A Scalable Framework for Optimizing Data Cleaning Pipelines for Machine Learning ApplicationsShafaq Siddiqi, Roman Kern, Matthias BoehmSIGMOD 2024 · 被引用 24 次
- CAPS: Cost-Aware ML Pipeline SelectionAntonios Kontaxakis, Dimitris Sacharidis, Alberto Abelló, Sergi Nadal 等VLDB 2026
- CleanML: A Study for Evaluating the Impact of Data Cleaning on ML Classification TasksPeng Li, Xi Rao, Jennifer Blase, Yue Zhang 等ICDE 2021 · 被引用 127 次
- SAPIENTML: Synthesizing Machine Learning Pipelines by Learning from Human-Written SolutionsRipon K. Saha, Akira Ura, Sonal Mahajan, Chenguang Zhu 等ICSE 2022 · 被引用 11 次
