HAIPipe: Combining Human-generated and Machine-generated Pipelines for Data Preparation
Sibei Chen, Nan Tang, Ju Fan, Xuemi Yan, Chengliang Chai, Guoliang Li, Xiaoyong Du
摘要
Data preparation is crucial in achieving optimized results for machine learning (ML). However, having a good data preparation pipeline is highly non-trivial for ML practitioners, which is not only domain-specific, but also dataset-specific. There are two common practices. Human-generated pipelines (HI-pipelines) typically use a wide range of any operations or libraries but are highly experience- and heuristic-based. In contrast, machine-generated pipelines (AI-pipelines), a.k.a. AutoML, often adopt a predefined set of sophisticated operations and are search-based and optimized. These two common practices are mutually complementary. In this paper, we study a new problem that, given an HI-pipeline and an AI-pipeline for the same ML task, can we combine them to get a new pipeline (HAI-pipeline) that is better than the provided HI-pipeline and AI-pipeline? We propose HAIPipe, a framework to address the problem, which adopts an enumeration-sampling strategy to carefully select the best performing combined pipeline. We also introduce a reinforcement learning (RL) based approach to search an optimized AI-pipeline. Extensive experiments using 1400+ real-world HI-pipelines (Jupyter notebooks from Kaggle) verify that HAIPipe can significantly outperform the approaches using either HI-pipelines or AI-pipelines alone.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- In-depth Analysis of Graph-based RAG in a Unified FrameworkYingli Zhou, Yaodong Su, Youran Sun, Shu Wang 等VLDB 2025 · 被引用 48 次
- ArchRAG: Attributed Community-based Hierarchical Retrieval-Augmented GenerationShu Wang, Yixiang Fang, Yingli Zhou, Xilin Liu 等AAAI 2026 · 被引用 23 次
- BookRAG: A Hierarchical Structure-aware Index-based Approach for Retrieval-Augmented Generation on Complex DocumentsShu Wang, Yingli Zhou, Yixiang FangVLDB 2026 · 被引用 16 次
- Automatic Database Configuration Debugging using Retrieval-Augmented Language ModelsSibei Chen, Ju Fan, Bin Wu, Nan Tang 等SIGMOD 2025 · 被引用 12 次
- AutoPrep: Natural Language Question-Aware Data Preparation with a Multi-Agent FrameworkMeihao Fan, Ju Fan, Nan Tang, Lei Cao 等VLDB 2025 · 被引用 10 次
它引用的顶会 Paper7
- Auto-Suggest: Learning-to-Recommend Data Preparation Steps Using Data Science NotebooksCong Yan, Yeye HeSIGMOD 2020 · 被引用 64 次
- Domain Adaptation for Deep Entity ResolutionJianhong Tu, Ju Fan, Nan Tang, Peng Wang 等SIGMOD 2022 · 被引用 46 次
- Vamsa: Automated Provenance Tracking in Data Science ScriptsMohammad Hossein Namaki, Avrilia Floratou, Fotis Psallidas, Subru Krishnan 等KDD 2020 · 被引用 41 次
- Semantics Driven Embedding Learning for Effective Entity AlignmentZiyue Zhong, Meihui Zhang, Ju Fan, Chenxiao DouICDE 2022 · 被引用 38 次
- Auto-Pipeline: Synthesize Data Pipelines By-Target Using Reinforcement Learning and SearchJunwen Yang, Yeye He, Surajit ChaudhuriVLDB 2021 · 被引用 32 次
相关 Paper
- DeepLine: AutoML Tool for Pipelines Generation using Deep Reinforcement Learning and Hierarchical Actions FilteringYuval Heffetz, Roman Vainshtein, Gilad Katz, Lior RokachKDD 2020 · 被引用 3 次
- CtxPipe: Context-aware Data Preparation Pipeline Construction for Machine LearningHaotian Gao, Shaofeng Cai, Tien Tuan Anh Dinh, Zhiyong Huang 等SIGMOD 2025 · 被引用 7 次
- SAPIENTML: Synthesizing Machine Learning Pipelines by Learning from Human-Written SolutionsRipon K. Saha, Akira Ura, Sonal Mahajan, Chenguang Zhu 等ICSE 2022 · 被引用 11 次
- Searching for Machine Learning Pipelines Using a Context-Free GrammarRadu Marinescu, Akihiro Kishimoto, Parikshit Ram, Ambrish Rawat 等AAAI 2021 · 被引用 18 次
- DiffPrep: Differentiable Data Preprocessing Pipeline Search for Learning over Tabular DataPeng Li, Zhiyi Chen, Xu Chu, Kexin RongSIGMOD 2023 · 被引用 24 次
