Auto-Suggest: Learning-to-Recommend Data Preparation Steps Using Data Science Notebooks
Cong Yan, Yeye He
摘要
Data preparation is widely recognized as the most timeconsuming process in modern business intelligence (BI) and machine learning (ML) projects. Automating complex data preparation steps (e.g., Pivot, Unpivot, Normalize-JSON, etc.) holds the potential to greatly improve user productivity, and has therefore become a central focus of research.
We propose a novel approach to "auto-suggest" contextualized data preparation steps, by "learning" from how data scientists would manipulate data, which are documented by data science notebooks widely available today. Specifically, we crawled over 4M Jupyter notebooks on GitHub, and replayed them step-by-step, to observe not only full input/output tables (data-frames) at each step, but also the exact data-preparation choices data scientists make that they believe are best suited to the input data (e.g., how input tables are Joined/Pivoted/Unpivoted, etc.). 1 By essentially "logging" how data scientists interact with diverse tables, and using the resulting logs as a proxy of "ground truth", we can learn-to-recommend data preparation steps best suited to given user data, just like how search engines (Google or Bing) leverage their click-through logs to learn-to-rank documents. This data-driven and log-driven approach leverages the "collective wisdom" of data scientists embodied in the notebooks, and is shown to significantly outperform strong baselines including commercial systems in terms of accuracy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper27
- Lux: Always-on Visualization Recommendations for Exploratory Dataframe WorkflowsDoris Jung Lin Lee, Dixin Tang, Kunal Agarwal, Thyne Boonmark 等VLDB 2022 · 被引用 61 次
- Improving Steering and Verification in AI-Assisted Data Analysis with Interactive Task DecompositionMajeed Kazemitabaar, Jack Williams, Ian Drosos, Tovi Grossman 等UIST 2024 · 被引用 49 次
- Fine-Grained Lineage for Safer Notebook InteractionsStephen Macke, Aditya G. Parameswaran, Hongpu Gong, Doris Jung Lin Lee 等VLDB 2021 · 被引用 46 次
- Data Formulator: AI-Powered Concept-Driven Visualization AuthoringChenglong Wang, John Thompson, Bongshin LeeIEEE VIS 2023 · 被引用 35 次
- CHORUS: Foundation Models for Unified Data Discovery and ExplorationMoe Kayali, Anton Lykov, Ilias Fountalis, Nikolaos Vasiloglou 等VLDB 2024 · 被引用 33 次
它引用的顶会 Paper1
相关 Paper
- Auto-Prep: Holistic Prediction of Data Preparation Steps for Self-Service Business IntelligenceEugenie Lai, Yeye He, Surajit ChaudhuriVLDB 2025 · 被引用 10 次
- Auto-Pipeline: Synthesize Data Pipelines By-Target Using Reinforcement Learning and SearchJunwen Yang, Yeye He, Surajit ChaudhuriVLDB 2021 · 被引用 32 次
- DiffPrep: Differentiable Data Preprocessing Pipeline Search for Learning over Tabular DataPeng Li, Zhiyi Chen, Xu Chu, Kexin RongSIGMOD 2023 · 被引用 24 次
- Subtle Bugs Everywhere: Generating Documentation for Data Wrangling CodeChenyang Yang, Shurui Zhou, Jin L. C. Guo, Christian KästnerASE 2021 · 被引用 25 次
- Cell2Doc: ML Pipeline for Generating Documentation in Computational NotebooksTamal Mondal, Scott Barnett, Akash Lal, Jyothi VeduradaASE 2023 · 被引用 4 次
