DiffPrep: Differentiable Data Preprocessing Pipeline Search for Learning over Tabular Data
Peng Li, Zhiyi Chen, Xu Chu, Kexin Rong
摘要
Data preprocessing is a crucial step in the machine learning process that transforms raw data into a more usable format for downstream ML models. However, it can be costly and time-consuming, often requiring the expertise of domain experts. Existing automated machine learning (AutoML) frameworks claim to automate data preprocessing. However, they often use a restricted search space of data preprocessing pipelines which limits the potential performance gains, and they are often too slow as they require training the ML model multiple times. In this paper, we propose DiffPrep, a method that can automatically and efficiently search for a data preprocessing pipeline for a given tabular dataset and a differentiable ML model such that the performance of the ML model is maximized. We formalize the problem of data preprocessing pipeline search as a bi-level optimization problem. To solve this problem efficiently, we transform and relax the discrete, non-differential search space into a continuous and differentiable one, which allows us to perform the pipeline search using gradient descent with training the ML model only once. Our experiments show that DiffPrep achieves the best test accuracy on 15 out of the 18 real-world datasets evaluated and improves the model's test accuracy by up to 6.6 percentage points.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Stress-Testing ML Pipelines with Adversarial Data CorruptionJiongli Zhu, Geyang Xu, Felipe Lorenzi, Boris Glavic 等VLDB 2025 · 被引用 2 次
- LLM-AutoDP: Automatic Data Processing via LLM Agents for Model Fine-tuningWei Huang, Anda Cheng, Yinggui Wang, Lei Wang 等VLDB 2026 · 被引用 1 次
它引用的顶会 Paper5
- Creating Embeddings of Heterogeneous Relational Datasets for Data Integration TasksRiccardo Cappuzzo, Paolo Papotti, Saravanan ThirumuruganathanSIGMOD 2020 · 被引用 139 次
- CleanML: A Study for Evaluating the Impact of Data Cleaning on ML Classification TasksPeng Li, Xi Rao, Jennifer Blase, Yue Zhang 等ICDE 2021 · 被引用 127 次
- Nearest Neighbor Classifiers over Incomplete Information: From Certain Answers to Certain PredictionsBojan Karlas, Peng Li, Renzhi Wu, Nezihe Merve Gürel 等VLDB 2021 · 被引用 69 次
- Evolving Search Space for Neural Architecture SearchYuanzheng Ci, Chen Lin, Ming Sun, Boyu Chen 等ICCV 2021 · 被引用 48 次
- Leva: Boosting Machine Learning Performance with Relational Embedding Data AugmentationZixuan Zhao, Raul Castro FernandezSIGMOD 2022 · 被引用 19 次
相关 Paper
- DAGPipe: Differentiable DAG Learning for Automated Data PreparationJing Chang, Chang LiuKDD 2026
- CtxPipe: Context-aware Data Preparation Pipeline Construction for Machine LearningHaotian Gao, Shaofeng Cai, Tien Tuan Anh Dinh, Zhiyong Huang 等SIGMOD 2025 · 被引用 7 次
- HAIPipe: Combining Human-generated and Machine-generated Pipelines for Data PreparationSibei Chen, Nan Tang, Ju Fan, Xuemi Yan 等SIGMOD 2023 · 被引用 25 次
- Auto-Suggest: Learning-to-Recommend Data Preparation Steps Using Data Science NotebooksCong Yan, Yeye HeSIGMOD 2020 · 被引用 64 次
- DMCO: Budget-Aware Co-Optimization of Data Cleaning and AutoMLXiaoou Ding, Zekai Qian, Siying Chen, Hongbin Hu 等ICML 2026
