DiffPrep: Differentiable Data Preprocessing Pipeline Search for Learning over Tabular Data
Peng Li, Zhiyi Chen, Xu Chu, Kexin Rong
Abstract
Data preprocessing is a crucial step in the machine learning process that transforms raw data into a more usable format for downstream ML models. However, it can be costly and time-consuming, often requiring the expertise of domain experts. Existing automated machine learning (AutoML) frameworks claim to automate data preprocessing. However, they often use a restricted search space of data preprocessing pipelines which limits the potential performance gains, and they are often too slow as they require training the ML model multiple times. In this paper, we propose DiffPrep, a method that can automatically and efficiently search for a data preprocessing pipeline for a given tabular dataset and a differentiable ML model such that the performance of the ML model is maximized. We formalize the problem of data preprocessing pipeline search as a bi-level optimization problem. To solve this problem efficiently, we transform and relax the discrete, non-differential search space into a continuous and differentiable one, which allows us to perform the pipeline search using gradient descent with training the ML model only once. Our experiments show that DiffPrep achieves the best test accuracy on 15 out of the 18 real-world datasets evaluated and improves the model's test accuracy by up to 6.6 percentage points.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 468584a3-d375-455d-8ff7-6575a499ce52Cited by top-tier papers2
- Stress-Testing ML Pipelines with Adversarial Data CorruptionJiongli Zhu, Geyang Xu, Felipe Lorenzi, Boris Glavic et al.VLDB 2025 · 2 citations
- LLM-AutoDP: Automatic Data Processing via LLM Agents for Model Fine-tuningWei Huang, Anda Cheng, Yinggui Wang, Lei Wang et al.VLDB 2026 · 1 citation
Builds on5
- Creating Embeddings of Heterogeneous Relational Datasets for Data Integration TasksRiccardo Cappuzzo, Paolo Papotti, Saravanan ThirumuruganathanSIGMOD 2020 · 139 citations
- CleanML: A Study for Evaluating the Impact of Data Cleaning on ML Classification TasksPeng Li, Xi Rao, Jennifer Blase, Yue Zhang et al.ICDE 2021 · 127 citations
- Nearest Neighbor Classifiers over Incomplete Information: From Certain Answers to Certain PredictionsBojan Karlas, Peng Li, Renzhi Wu, Nezihe Merve Gürel et al.VLDB 2021 · 69 citations
- Evolving Search Space for Neural Architecture SearchYuanzheng Ci, Chen Lin, Ming Sun, Boyu Chen et al.ICCV 2021 · 48 citations
- Leva: Boosting Machine Learning Performance with Relational Embedding Data AugmentationZixuan Zhao, Raul Castro FernandezSIGMOD 2022 · 19 citations
Related papers
- DAGPipe: Differentiable DAG Learning for Automated Data PreparationJing Chang, Chang LiuKDD 2026
- CtxPipe: Context-aware Data Preparation Pipeline Construction for Machine LearningHaotian Gao, Shaofeng Cai, Tien Tuan Anh Dinh, Zhiyong Huang et al.SIGMOD 2025 · 7 citations
- HAIPipe: Combining Human-generated and Machine-generated Pipelines for Data PreparationSibei Chen, Nan Tang, Ju Fan, Xuemi Yan et al.SIGMOD 2023 · 25 citations
- Auto-Suggest: Learning-to-Recommend Data Preparation Steps Using Data Science NotebooksCong Yan, Yeye HeSIGMOD 2020 · 64 citations
- DMCO: Budget-Aware Co-Optimization of Data Cleaning and AutoMLXiaoou Ding, Zekai Qian, Siying Chen, Hongbin Hu et al.ICML 2026
