CHORUS: Foundation Models for Unified Data Discovery and Exploration
Moe Kayali, Anton Lykov, Ilias Fountalis, Nikolaos Vasiloglou, Dan Olteanu, Dan Suciu
Abstract
We apply foundation models to data discovery and exploration tasks. Foundation models are large language models (llms) that show promising performance on a range of diverse tasks unrelated to their training. We show that these models are highly applicable to the data discovery and data exploration domain. When carefully used, they have superior capability on three representative tasks: table-class detection, column-type annotation and join-column prediction. On all three tasks, we show that a foundation-model-based approach outperforms the task-specific models and so the state of the art. Further, our approach often surpasses human-expert task performance. We investigate the fundamental characteristics of this approach including generalizability to several foundation models and the impact of non-determinism on the outputs. All in all, this suggests a future direction in which disparate data management tasks can be unified under foundation models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c94ceafb-fa62-48d0-b385-a1e19b0c2a84Cited by top-tier papers22
- Table-GPT: Table Fine-tuned GPT for Diverse Table TasksPeng Li, Yeye He, Dror Yashar, Weiwei Cui et al.SIGMOD 2024 · 63 citations
- DocETL: Agentic Query Rewriting and Evaluation for Complex Document ProcessingShreya Shankar, Tristan Chambers, Tarak Shah, Aditya G. Parameswaran et al.VLDB 2025 · 62 citations
- Cost-Effective In-Context Learning for Entity Resolution: A Design Space ExplorationMeihao Fan, Xiaoyue Han, Ju Fan, Chengliang Chai et al.ICDE 2024 · 40 citations
- ArcheType: A Novel Framework for Open-Source Column Type Annotation using Large Language ModelsBenjamin Feuer, Yurong Liu, Chinmay Hegde, Juliana FreireVLDB 2024 · 33 citations
- Magneto: Combining Small and Large Language Models for Schema MatchingYurong Liu, Eduardo H. M. Pena, Aécio S. R. Santos, Eden Wu et al.VLDB 2025 · 32 citations
Builds on19
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
- TURL: Table Understanding through Representation LearningXiang Deng, Huan Sun, Alyssa Lees, You Wu et al.VLDB 2021 · 2,406 citations
Related papers
- Can Foundation Models Wrangle Your Data?Avanika Narayan, Ines Chami, Laurel J. Orr, Christopher RéVLDB 2023 · 325 citations
- Unveiling Challenges for LLMs in Enterprise Data EngineeringJan-Micha Bodensohn, Ulf Brackmann, Liane Vogel, Anupam Sanghi et al.VLDB 2026 · 13 citations
- Fundamental Capabilities of Large Language Models and their Applications in Domain Scenarios: A SurveyJiawei Li, Yizhe Yang, Yu Bai, Xiaofeng Zhou et al.ACL 2024 · 15 citations
- Text2Analysis: A Benchmark of Table Question Answering with Advanced Data Analysis and Unclear QueriesXinyi He, Mengyu Zhou, Xinrun Xu, Xiaojun Ma et al.AAAI 2024 · 48 citations
- Label Annotation for Tabular Anomaly Detection with Large Language ModelsHaihong Zhao, Aochuan Chen, Miao Peng, Xiaolong Fan et al.KDD 2026
