QueryArtisan: Generating Data Manipulation Codes for Ad-hoc Analysis in Data Lakes
Xiu Tang, Wenhao Liu, Sai Wu, Chang Yao, Gongsheng Yuan, Shanshan Ying, Gang Chen
摘要
Query processing over data lakes is a challenging task, often requiring extensive data pre-processing activities such as data cleaning, transformation, and loading. However, the advent of Large Language Models (LLMs) has illuminated a new pathway to address these complexities by offering a unified approach to understanding the diverse datasets submerged in data lakes. In this paper, we introduce QueryArtisan, a novel LLM-powered analytic tool specifically designed for data lakes. QueryArtisan transcends traditional ETL (Extract, Transform, Load) processes by generating just-intime code for dataset-specific queries. It eliminates the need for an intermediary schema, enabling users to query the data lake directly using natural language. To achieve this, we have developed a suite of heterogeneous operators capable of processing data across various modalities. Additionally, QueryArtisan incorporates a cost model-based query optimization technique, significantly enhancing its code generation capabilities for efficient query resolution. Our extensive experimental evaluations, conducted with real-life datasets, demonstrate that QueryArtisan markedly outperforms existing solutions in terms of effectiveness, efficiency and usability.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper19
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu 等NeurIPS 2023 · 被引用 5,989 次
- TURL: Table Understanding through Representation LearningXiang Deng, Huan Sun, Alyssa Lees, You Wu 等VLDB 2021 · 被引用 2,406 次
- Text-to-SQL Empowered by Large Language Models: A Benchmark EvaluationDawei Gao, Haibin Wang, Yaliang Li, Xiuyu Sun 等VLDB 2024 · 被引用 609 次
- TAPEX: Table Pre-training via Learning a Neural SQL ExecutorQian Liu, Bei Chen, Jiaqi Guo, Morteza Ziyadi 等ICLR 2022 · 被引用 347 次
相关 Paper
- Logical and Physical Optimizations for SQL Query Execution over Large Language ModelsDario Satriani, Enzo Veltri, Donatello Santoro, Sara Rosato 等SIGMOD 2025 · 被引用 7 次
- SEMA: A High-performance System for LLM-based Semantic Query ProcessingKangkang Qi, Dongyang Xie, Wenbo Li, Hao Zhang 等VLDB 2026 · 被引用 5 次
- Can Large Language Models Be Query Optimizer for Relational Databases?Jie Tan, Kangfei Zhao, Rui Li, Jeffrey Xu Yu 等SIGMOD 2026 · 被引用 6 次
- UQE: A Query Engine for Unstructured DatabasesHanjun Dai, Bethany Wang, Xingchen Wan, Bo Dai 等NeurIPS 2024 · 被引用 45 次
- ViseGPT: Towards Better Alignment of LLM-generated Data Wrangling Scripts and User PromptsJiajun Zhu, Xinyu Cheng, Zhongsu Luo, Yunfan Zhou 等UIST 2025 · 被引用 1 次
