EcoTable: Cost-effective Table Integration in Data Lakes for Natural Language Queries
Yuhui Wang, Jinqi Liu, Chengliang Chai, Hangyu Zhao, Yuhao Deng, Yuyu Luo, Xin Tang, Ye Yuan, Guoren Wang, Fengjin Wang, Lei Cao
摘要
The diverse formats of CSV and Parquet files in data lakes pose a significant challenge to traditional ETL, which relies on data engineers to pre-define a target database schema and build a complex pipeline for data integration. Moreover, with this approach, the integrated data often cannot support various analytical needs, as the predefined schema does not necessarily satisfy the table format or join relationships required to answer unforeseen queries. To address this, we propose EcoTable , the first natural languagebased data integration framework. Given a set of user-specified natural language queries, EcoTable automatically integrates the tables into a form that adequately supports the corresponding SQL queries. EcoTable achieves this by leveraging the semantic understanding and complex reasoning capabilities of Large Language Models (LLMs). Moreover, EcoTable addresses the scalability and cost issues introduced by expensive LLM inferences with a set of novel ideas.
EcoTable first introduces a graph to represent the overall search space, where nodes represent tables and edges carry weights indicating join likelihood produced by a lightweight deep learning model. On top of this graph data structure, EcoTable designs three components to achieve our goal: (1) the table identification layer aims to identify relevant tables via a two-stage schema linking based on user queries; (2) the graph-based validation layer aims to discover significant join paths, including necessary data transformations and bridging tables, by modeling the problem as Steiner tree searches; and (3) the table transformation layer generates transformation code to implement the joins using LLMs. We construct * Chengliang Chai is the corresponding author.
4 real-world benchmark datasets with more than 200 queries. Extensive experiments demonstrate that EcoTable outperforms the state-of-the-art baselines, increasing accuracy by more than 30% and cutting LLM invocation costs by 5 times.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper27
- DIN-SQL: Decomposed In-Context Learning of Text-to-SQL with Self-CorrectionMohammadreza Pourreza, Davood RafieiNeurIPS 2023 · 被引用 909 次
- Deep Entity Matching with Pre-Trained Language ModelsYuliang Li, Jinfeng Li, Yoshihiko Suhara, AnHai Doan 等VLDB 2021 · 被引用 484 次
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding SharingPengcheng He, Jianfeng Gao, Weizhu ChenICLR 2023 · 被引用 394 次
- RESDSQL: Decoupling Schema Linking and Skeleton Parsing for Text-to-SQLHaoyang Li, Jing Zhang, Cuiping Li, Hong ChenAAAI 2023 · 被引用 343 次
- Plan-and-Solve Prompting: Improving Zero-Shot Chain-of-Thought Reasoning by Large Language ModelsLei Wang, Wanyu Xu, Yihuai Lan, Zhiqiang Hu 等ACL 2023 · 被引用 249 次
相关 Paper
- Weaver: Interweaving SQL and LLM for Table ReasoningRohit Khoja, Devanshu Gupta, Yanjie Fu, Dan Roth 等EMNLP 2025 · 被引用 1 次
- ReAcTable: Enhancing ReAct for Table Question AnsweringYunjia Zhang, Jordan Henkel, Avrilia Floratou, Joyce Cahoon 等VLDB 2024 · 被引用 120 次
- DeepJoin: Joinable Table Discovery with Pre-trained Language ModelsYuyang Dong, Chuan Xiao, Takuma Nozawa, Masafumi Enomoto 等VLDB 2023 · 被引用 53 次
- QueryArtisan: Generating Data Manipulation Codes for Ad-hoc Analysis in Data LakesXiu Tang, Wenhao Liu, Sai Wu, Chang Yao 等VLDB 2025 · 被引用 4 次
- GRIT: Guided Relational Integration for Efficient Multi-Table UnderstandingYujin Kang, Park Seong Woo, Yoon-Sik ChoEMNLP 2025 · 被引用 2 次
