Auto-BI: Automatically Build BI-Models Leveraging Local Join Prediction and Global Schema Graph
Yiming Lin, Yeye He, Surajit Chaudhuri
Abstract
Business Intelligence (BI) is crucial in modern enterprises and billion-dollar business. Traditionally, technical experts like database administrators would manually prepare BI-models (e.g., in star or snowflake schemas) that join tables in data warehouses, before less-technical business users can run analytics using end-user dashboarding tools. However, the popularity of self-service BI (e.g., Tableau and Power-BI) in recent years creates a strong demand for less technical end-users to build BI-models themselves.
We develop an Auto-BI system that can accurately predict BI models given a set of input tables, using a principled graph-based optimization problem we propose called k-Min-Cost-Arborescence (k-MCA), which holistically considers both local join prediction and global schema-graph structures, leveraging a graph-theoretical structure called arborescence. While we prove k-MCA is intractable and inapproximate in general, we develop novel algorithms that can solve k-MCA optimally, which is shown to be efficient in practice with sub-second latency and can scale to the largest BI-models we encounter (with close to 100 tables).
Auto-BI is rigorously evaluated on a unique dataset with over 100K real BI models we harvested, as well as on 4 popular TPC benchmarks. It is shown to be both efficient and accurate, achieving over 0.9 F1-score on both real and synthetic benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Auto-Tables: Synthesizing Multi-Step Transformations to Relationalize Tables without Using ExamplesPeng Li, Yeye He, Cong Yan, Yue Wang et al.VLDB 2023 · 29 citations
- Auto-Prep: Holistic Prediction of Data Preparation Steps for Self-Service Business IntelligenceEugenie Lai, Yeye He, Surajit ChaudhuriVLDB 2025 · 10 citations
- BIRDIE: Natural Language-Driven Table Discovery Using Differentiable Search IndexYuxiang Guo, Zhonghao Hu, Yuren Mao, Baihua Zheng et al.VLDB 2025 · 6 citations
- Auto-Test: Learning Semantic-Domain Constraints for Unsupervised Error Detection in TablesQixu Chen, Yeye He, Raymond Chi-Wing Wong, Weiwei Cui et al.SIGMOD 2025 · 4 citations
- Auto-Fill: Learning to Predict Missing Values Accurately with Specialist Language ModelsYurong Liu, Yeye He, Haoyu Dong, Junjie Xing et al.VLDB 2026
Builds on4
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Auto-Suggest: Learning-to-Recommend Data Preparation Steps Using Data Science NotebooksCong Yan, Yeye HeSIGMOD 2020 · 64 citations
- Proton: Probing Schema Linking Information from Pre-trained Language Models for Text-to-SQL ParsingLihan Wang, Bowen Qin, Binyuan Hui, Bowen Li et al.KDD 2022 · 31 citations
- Efficiently Transforming Tables for JoinabilityArash Dargahi Nobari, Davood RafieiICDE 2022 · 8 citations
Related papers
- Scalable and Usable Relational Learning With Automatic Language BiasJose Picado, Arash Termehchy, Alan Fern, Sudhanshu Pathak et al.SIGMOD 2021 · 5 citations
- AutoCE: An Accurate and Efficient Model Advisor for Learned Cardinality EstimationJintao Zhang, Chao Zhang, Guoliang Li, Chengliang ChaiICDE 2023 · 13 citations
- DashBot: Insight-Driven Dashboard Generation Based on Deep Reinforcement LearningDazhen Deng, Aoyu Wu, Huamin Qu, Yingcai WuIEEE VIS 2022 · 40 citations
- Table2Analysis: Modeling and Recommendation of Common Analysis Patterns for Multi-Dimensional DataMengyu Zhou, Wang Tao, Pengxin Ji, Han Shi et al.AAAI 2020 · 26 citations
- Auto-Formula: Recommend Formulas in Spreadsheets using Contrastive Learning for Table RepresentationsSibei Chen, Yeye He, Weiwei Cui, Ju Fan et al.SIGMOD 2024 · 4 citations
