Can Large Language Models Predict Data Correlations from Column Names?
Immanuel Trummer
摘要
Recent publications suggest using natural language analysis on database schema elements to guide tuning and profiling efforts. The underlying hypothesis is that state-of-the-art language processing methods, so-called language models, are able to extract information on data properties from schema text. This paper examines that hypothesis in the context of data correlation analysis: is it possible to find column pairs with correlated data by analyzing their names via language models? First, the paper introduces a novel benchmark for data correlation analysis, created by analyzing thousands of Kaggle data sets (and available for download). Second, it uses that data to study the ability of language models to predict correlation, based on column names. The analysis covers different language models, various correlation metrics, and a multitude of accuracy metrics. It pinpoints factors that contribute to successful predictions, such as the length of column names as well as the ratio of words. Finally, the study analyzes the impact of column types on prediction performance. The results show that schema text can be a useful source of information and inform future research efforts, targeted at NLP-enhanced database tuning and data profiling.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- E2ETune: End-to-End Knob Tuning via Fine-tuned Generative Language ModelXinmei Huang, Haoyang Li, Jing Zhang, Xinxin Zhao 等VLDB 2025 · 被引用 15 次
- Chameleon: Foundation Models for Fairness-aware Multi-modal Data Augmentation to Enhance Coverage of MinoritiesMahdi Erfanian, H. V. Jagadish, Abolfazl AsudehVLDB 2024 · 被引用 10 次
- Exploiting LLMs for Automatic Hypothesis Assessment via a Logit-Based Calibrated PriorYue Gong, Raul Castro FernandezNeurIPS 2025 · 被引用 1 次
- Accurate Table Question Answering with Accessible LLMsYangfan Jiang, Fei Wei, Ergute Bao, Yaliang Li 等ICDE 2026 · 被引用 1 次
它引用的顶会 Paper15
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- Can Foundation Models Wrangle Your Data?Avanika Narayan, Ines Chami, Laurel J. Orr, Christopher RéVLDB 2023 · 被引用 325 次
- Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data LakesSimran Arora, Brandon Yang, Sabri Eyuboglu, Avanika Narayan 等VLDB 2024 · 被引用 165 次
- RPT: Relational Pre-trained Transformer Is Almost All You Need towards Democratizing Data PreparationNan Tang, Ju Fan, Fangyi Li, Jianhong Tu 等VLDB 2021 · 被引用 92 次
- CodexDB: Synthesizing Code for Query Processing from Natural Language Instructions using GPT-3 CodexImmanuel TrummerVLDB 2022 · 被引用 77 次
相关 Paper
- SchemaPile: A Large Collection of Relational Database SchemasTill Döhmen, Radu Geacu, Madelon Hulsebos, Sebastian SchelterSIGMOD 2024 · 被引用 9 次
- NameGuess: Column Name Expansion for Tabular DataJiani Zhang, Zhengyuan Shen, Balasubramaniam Srinivasan, Shen Wang 等EMNLP 2023 · 被引用 6 次
- The Case for NLP-Enhanced Database Tuning: Towards Tuning Tools that "Read the Manual"Immanuel TrummerVLDB 2021 · 被引用 27 次
- DB-BERT: A Database Tuning Tool that "Reads the Manual"Immanuel TrummerSIGMOD 2022 · 被引用 71 次
- Annotating Columns with Pre-trained Language ModelsYoshihiko Suhara, Jinfeng Li, Yuliang Li, Dan Zhang 等SIGMOD 2022 · 被引用 81 次
