Auto-Pipeline: Synthesize Data Pipelines By-Target Using Reinforcement Learning and Search
Junwen Yang, Yeye He, Surajit Chaudhuri
摘要
Recent work has made significant progress in helping users to automate single data preparation steps, such as string-transformations and table-manipulation operators (e.g., Join, GroupBy, Pivot, etc.). We in this work propose to automate multiple such steps end-to-end, by synthesizing complex data-pipelines with both string-transformations and table-manipulation operators.
We propose a novel by-target paradigm that allows users to easily specify the desired pipeline, which is a significant departure from the traditional by-example paradigm. Using by-target, users would provide input tables (e.g., csv or json files), and point us to a "target table" (e.g., an existing database table or BI dashboard) to demonstrate how the output from the desired pipeline would schematically "look like". While the problem is seemingly under-specified, our unique insight is that implicit table constraints such as FDs and keys can be exploited to significantly constrain the space and make the problem tractable. We develop an AUTO-PIPELINE system that learns to synthesize pipelines using deep reinforcement-learning (DRL) and search. Experiments using a benchmark of 700 real pipelines crawled from GitHub and commercial vendors suggest that AUTO-PIPELINE can successfully synthesize around 70% of complex pipelines with up to 10 steps.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- HAIPipe: Combining Human-generated and Machine-generated Pipelines for Data PreparationSibei Chen, Nan Tang, Ju Fan, Xuemi Yan 等SIGMOD 2023 · 被引用 25 次
- Rigel: Transforming Tabular Data by Declarative MappingRan Chen, Di Weng, Yanwei Huang, Xinhuan Shu 等IEEE VIS 2022 · 被引用 21 次
- Gen-T: Table Reclamation in Data LakesGrace Fan, Roee Shraga, Renée J. MillerICDE 2024 · 被引用 5 次
- Auto-Formula: Recommend Formulas in Spreadsheets using Contrastive Learning for Table RepresentationsSibei Chen, Yeye He, Weiwei Cui, Ju Fan 等SIGMOD 2024 · 被引用 4 次
- Ferry: Toward Better Understanding of Input/Output Space for Data Wrangling ScriptsZhongsu Luo, Kai Xiong, Jiajun Zhu, Ran Chen 等IEEE VIS 2024 · 被引用 4 次
它引用的顶会 Paper6
- Deep Entity Matching with Pre-Trained Language ModelsYuliang Li, Jinfeng Li, Yoshihiko Suhara, AnHai Doan 等VLDB 2021 · 被引用 484 次
- Auto-Suggest: Learning-to-Recommend Data Preparation Steps Using Data Science NotebooksCong Yan, Yeye HeSIGMOD 2020 · 被引用 64 次
- SCODED: Statistical Constraint Oriented Data Error DetectionJing Nathan Yan, Oliver Schulte, Mohan Zhang, Jiannan Wang 等SIGMOD 2020 · 被引用 32 次
- Searching a Database of Source Codes Using Contextualized Code SearchRohan Mukherjee, Chris Jermaine, Swarat ChaudhuriVLDB 2020 · 被引用 11 次
- Baran: Effective Error Correction via a Unified Context Representation and Transfer LearningMohammad Mahdavi, Ziawasch AbedjanVLDB 2020
相关 Paper
- DeepLine: AutoML Tool for Pipelines Generation using Deep Reinforcement Learning and Hierarchical Actions FilteringYuval Heffetz, Roman Vainshtein, Gilad Katz, Lior RokachKDD 2020 · 被引用 3 次
- Auto-Prep: Holistic Prediction of Data Preparation Steps for Self-Service Business IntelligenceEugenie Lai, Yeye He, Surajit ChaudhuriVLDB 2025 · 被引用 10 次
- Auto-Tables: Synthesizing Multi-Step Transformations to Relationalize Tables without Using ExamplesPeng Li, Yeye He, Cong Yan, Yue Wang 等VLDB 2023 · 被引用 29 次
- BAT: Target-Instance-Free Data Preparation Synthesis via LLM-Driven Tree SearchCongcong Ge, Yachuan Liu, Yixuan Tang, Yifan Zhu 等SIGMOD 2026
- CtxPipe: Context-aware Data Preparation Pipeline Construction for Machine LearningHaotian Gao, Shaofeng Cai, Tien Tuan Anh Dinh, Zhiyong Huang 等SIGMOD 2025 · 被引用 7 次
