Empowering Tabular Data Preparation with Language Models: Why and How?
Mengshi Chen, Yuxiang Sun, Tengchao Li, Jianwei Wang, Kai Wang, Xuemin Lin, Ying Zhang, Wenjie Zhang
摘要
Data preparation is a critical step in enhancing the usability of tabular data and thus boosts downstream data-driven tasks. Traditional methods often face challenges in capturing the intricate relationships within tables and adapting to the tasks involved. Recent advances in Language Models (LMs), especially in Large Language Models (LLMs), offer new opportunities to automate and support tabular data preparation. However, why LMs suit tabular data preparation (i.e., how their capabilities match task demands) and how to use them effectively across phases still remain to be systematically explored. In this survey, we systematically analyze the role of LMs in enhancing tabular data preparation processes, focusing on four core phases: data acquisition, integration, cleaning, and transformation. For each phase, we present an integrated analysis of how LMs can be combined with other components for different preparation tasks, highlight key advancements, and outline prospective pipelines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- PrepBench: How Far Are We from Natural-Language-Driven Data Preparation?Jingzhe Xu, Rui Wang, Jiannan Wang, Guoliang LiVLDB 2026 · 被引用 3 次
- Rethink Representation Learning for Questionnaire DataGuanhua Ye, Jifeng He, Yan Li, Junping Du 等AAAI 2026
它引用的顶会 Paper20
- MPNet: Masked and Permuted Pre-training for Language UnderstandingKaitao Song, Xu Tan, Tao Qin, Jianfeng Lu 等NeurIPS 2020 · 被引用 1,957 次
- Deep Entity Matching with Pre-Trained Language ModelsYuliang Li, Jinfeng Li, Yoshihiko Suhara, AnHai Doan 等VLDB 2021 · 被引用 484 次
- Can Foundation Models Wrangle Your Data?Avanika Narayan, Ines Chami, Laurel J. Orr, Christopher RéVLDB 2023 · 被引用 325 次
- Semantics-aware Dataset Discovery from Data Lakes with Contextualized Column-based Representation LearningGrace Fan, Jin Wang, Yuliang Li, Dan Zhang 等VLDB 2023 · 被引用 139 次
- Dual-Objective Fine-Tuning of BERT for Entity MatchingRalph Peeters, Christian BizerVLDB 2021 · 被引用 71 次
相关 Paper
- Unveiling Challenges for LLMs in Enterprise Data EngineeringJan-Micha Bodensohn, Ulf Brackmann, Liane Vogel, Anupam Sanghi 等VLDB 2026 · 被引用 13 次
- AutoPrep: Natural Language Question-Aware Data Preparation with a Multi-Agent FrameworkMeihao Fan, Ju Fan, Nan Tang, Lei Cao 等VLDB 2025 · 被引用 10 次
- Table Question Answering in the Era of Large Language Models: A Comprehensive Survey of Tasks, Methods, and EvaluationWei Zhou, Bolei Ma, Annemarie Friedrich, Mohsen MesgarACL 2026 · 被引用 3 次
- Visualization Recommendation with Prompt-based Reprogramming of Large Language ModelsXinhang Li, Jingbo Zhou, Wei Chen, Derong Xu 等ACL 2024 · 被引用 4 次
- BAT: Target-Instance-Free Data Preparation Synthesis via LLM-Driven Tree SearchCongcong Ge, Yachuan Liu, Yixuan Tang, Yifan Zhu 等SIGMOD 2026
