Empowering Tabular Data Preparation with Language Models: Why and How?
Mengshi Chen, Yuxiang Sun, Tengchao Li, Jianwei Wang, Kai Wang, Xuemin Lin, Ying Zhang, Wenjie Zhang
Abstract
Data preparation is a critical step in enhancing the usability of tabular data and thus boosts downstream data-driven tasks. Traditional methods often face challenges in capturing the intricate relationships within tables and adapting to the tasks involved. Recent advances in Language Models (LMs), especially in Large Language Models (LLMs), offer new opportunities to automate and support tabular data preparation. However, why LMs suit tabular data preparation (i.e., how their capabilities match task demands) and how to use them effectively across phases still remain to be systematically explored. In this survey, we systematically analyze the role of LMs in enhancing tabular data preparation processes, focusing on four core phases: data acquisition, integration, cleaning, and transformation. For each phase, we present an integrated analysis of how LMs can be combined with other components for different preparation tasks, highlight key advancements, and outline prospective pipelines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ca637267-7418-488a-8fdf-398994b0de0eCited by top-tier papers2
- PrepBench: How Far Are We from Natural-Language-Driven Data Preparation?Jingzhe Xu, Rui Wang, Jiannan Wang, Guoliang LiVLDB 2026 · 3 citations
- Rethink Representation Learning for Questionnaire DataGuanhua Ye, Jifeng He, Yan Li, Junping Du et al.AAAI 2026
Builds on20
- MPNet: Masked and Permuted Pre-training for Language UnderstandingKaitao Song, Xu Tan, Tao Qin, Jianfeng Lu et al.NeurIPS 2020 · 1,957 citations
- Deep Entity Matching with Pre-Trained Language ModelsYuliang Li, Jinfeng Li, Yoshihiko Suhara, AnHai Doan et al.VLDB 2021 · 484 citations
- Can Foundation Models Wrangle Your Data?Avanika Narayan, Ines Chami, Laurel J. Orr, Christopher RéVLDB 2023 · 325 citations
- Semantics-aware Dataset Discovery from Data Lakes with Contextualized Column-based Representation LearningGrace Fan, Jin Wang, Yuliang Li, Dan Zhang et al.VLDB 2023 · 139 citations
- Dual-Objective Fine-Tuning of BERT for Entity MatchingRalph Peeters, Christian BizerVLDB 2021 · 71 citations
Related papers
- Unveiling Challenges for LLMs in Enterprise Data EngineeringJan-Micha Bodensohn, Ulf Brackmann, Liane Vogel, Anupam Sanghi et al.VLDB 2026 · 13 citations
- AutoPrep: Natural Language Question-Aware Data Preparation with a Multi-Agent FrameworkMeihao Fan, Ju Fan, Nan Tang, Lei Cao et al.VLDB 2025 · 10 citations
- Table Question Answering in the Era of Large Language Models: A Comprehensive Survey of Tasks, Methods, and EvaluationWei Zhou, Bolei Ma, Annemarie Friedrich, Mohsen MesgarACL 2026 · 3 citations
- Visualization Recommendation with Prompt-based Reprogramming of Large Language ModelsXinhang Li, Jingbo Zhou, Wei Chen, Derong Xu et al.ACL 2024 · 4 citations
- BAT: Target-Instance-Free Data Preparation Synthesis via LLM-Driven Tree SearchCongcong Ge, Yachuan Liu, Yixuan Tang, Yifan Zhu et al.SIGMOD 2026
