Auto-Prep: Holistic Prediction of Data Preparation Steps for Self-Service Business Intelligence
Eugenie Lai, Yeye He, Surajit Chaudhuri
Abstract
Business Intelligence (BI) plays a critical role in empowering modern enterprises to make informed data-driven decisions, and has grown into a billion-dollar business. Self-service BI tools like Power BI and Tableau have democratized the "dashboarding" phase of BI, by offering user-friendly, drag-and-drop interfaces that are tailored to non-technical enterprise users. However, despite these advances, we observe that the "data preparation" phase of BI continues to be a key pain point for BI users today. In this work, we systematically study around 2K real BI projects harvested from public sources, focusing on the data-preparation phase of the BI workflows. We observe that users often have to program both (1) data transformation steps and (2) table joins steps, before their raw data can be ready for dashboarding and analysis. A careful study of the BI workflows reveals that transformation and join steps are often intertwined in the same BI project, such that considering both holistically is crucial to accurately predict these steps. Leveraging this observation, we develop an Auto-Prep system to holistically predict transformations and joins, using a principled graph-based algorithm inspired by Steiner-tree, with provable quality guarantees. Extensive evaluations using real BI projects suggest that Auto-Prep can correctly predict over 70% transformation and join steps, significantly more accurate than existing algorithms as well as language-models such as GPT-4.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- DeepPrep: An LLM-Powered Agentic System for Autonomous Data PreparationMeihao Fan, Ju Fan, Yuxin Zhang, Shaolei Zhang et al.VLDB 2026 · 4 citations
- EcoTable: Cost-effective Table Integration in Data Lakes for Natural Language QueriesYuhui Wang, Jinqi Liu, Chengliang Chai, Hangyu Zhao et al.VLDB 2026
Builds on11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Can Foundation Models Wrangle Your Data?Avanika Narayan, Ines Chami, Laurel J. Orr, Christopher RéVLDB 2023 · 325 citations
- Auto-Suggest: Learning-to-Recommend Data Preparation Steps Using Data Science NotebooksCong Yan, Yeye HeSIGMOD 2020 · 64 citations
Related papers
- Auto-BI: Automatically Build BI-Models Leveraging Local Join Prediction and Global Schema GraphYiming Lin, Yeye He, Surajit ChaudhuriVLDB 2023 · 17 citations
- Auto-Pipeline: Synthesize Data Pipelines By-Target Using Reinforcement Learning and SearchJunwen Yang, Yeye He, Surajit ChaudhuriVLDB 2021 · 32 citations
- Auto-Tables: Synthesizing Multi-Step Transformations to Relationalize Tables without Using ExamplesPeng Li, Yeye He, Cong Yan, Yue Wang et al.VLDB 2023 · 29 citations
- Transitioning to a Commercial Dashboarding System: Socio-Technical Observations and OpportunitiesConny Walchshofer, Vaishali Dhanoa, Marc Streit, Miriah MeyerIEEE VIS 2023 · 15 citations
- Auto-Transform: Learning-to-Transform by PatternsYeye He, Zhongjun Jin, Surajit ChaudhuriVLDB 2020
