Unveiling Challenges for LLMs in Enterprise Data Engineering
Jan-Micha Bodensohn, Ulf Brackmann, Liane Vogel, Anupam Sanghi, Carsten Binnig
摘要
Large Language Models (LLMs) promise to automate data engineering on tabular data, offering enterprises a valuable opportunity to cut the high costs of manual data handling. But the enterprise domain comes with unique challenges that existing LLM-based approaches for data engineering often overlook, such as large table sizes, more complex tasks, and the need for internal knowledge. To bridge these gaps, we identify key enterprise-specific challenges related to data, tasks, and background knowledge and extensively evaluate how they affect data engineering with LLMs. Our analysis reveals that LLMs face substantial limitations in real-world enterprise scenarios, with accuracy declining sharply. Our findings contribute to a systematic understanding of LLMs for enterprise data engineering to support their adoption in industry.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- LLMs as World Models: Data-Driven and Human-Centered Pre-Event Simulation for Disaster Impact AssessmentLingyao Li, Dawei Li, Zhenhui Ou, Xiaoran Xu 等EMNLP 2025 · 被引用 5 次
- On Path to Multimodal Historical Reasoning: HistBench and HistAgentJiahao Qiu, Fulian Xiao, Yimin Wang, Yuchen Mao 等ICML 2026 · 被引用 5 次
- Mixtera: A Data Plane for Foundation Model TrainingMaximilian Böther, Xiaozhe Yao, Tolga Kerimoglu, Dan Graur 等SIGMOD 2026 · 被引用 5 次
它引用的顶会 Paper21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- TURL: Table Understanding through Representation LearningXiang Deng, Huan Sun, Alyssa Lees, You Wu 等VLDB 2021 · 被引用 2,406 次
- Text-to-SQL Empowered by Large Language Models: A Benchmark EvaluationDawei Gao, Haibin Wang, Yaliang Li, Xiuyu Sun 等VLDB 2024 · 被引用 609 次
相关 Paper
- Empowering Tabular Data Preparation with Language Models: Why and How?Mengshi Chen, Yuxiang Sun, Tengchao Li, Jianwei Wang 等ACL 2026 · 被引用 4 次
- CHORUS: Foundation Models for Unified Data Discovery and ExplorationMoe Kayali, Anton Lykov, Ilias Fountalis, Nikolaos Vasiloglou 等VLDB 2024 · 被引用 33 次
- Towards Synthetic Trace Generation of Modeling Operations using In-Context Learning ApproachVittoriano Muttillo, Claudio Di Sipio, Riccardo Rubei, Luca Berardinelli 等ASE 2024 · 被引用 1 次
- Human-LLM Collaborative Feature Engineering for Tabular DataZhuoyan Li, Aditya Bansal, Jinzhao Li, Shishuang He 等ICLR 2026 · 被引用 2 次
- Large Language Models for Automated Data Science: Introducing CAAFE for Context-Aware Automated Feature EngineeringNoah Hollmann, Samuel Müller, Frank HutterNeurIPS 2023 · 被引用 210 次
