Logical and Physical Optimizations for SQL Query Execution over Large Language Models
Dario Satriani, Enzo Veltri, Donatello Santoro, Sara Rosato, Simone Varriale, Paolo Papotti
摘要
Interacting with Large Language Models (LLMs) via declarative queries is increasingly popular for tasks like question answering and data extraction, thanks to their ability to process vast unstructured data. However, LLMs often struggle with answering complex factual questions, exhibiting low precision and recall in the returned data. This challenge highlights that executing queries on LLMs remains a largely unexplored domain, where traditional data processing assumptions often fall short. Conventional query optimization, typically cost-driven, overlooks LLM-specific quality challenges such as contextual understanding. Just as new physical operators are designed to address the unique characteristics of LLMs, optimization must consider these quality challenges. Our results highlight that adhering strictly to conventional query optimization principles fails to generate the best plans in terms of result quality. To tackle this challenge, we present a novel approach to enhance SQL results by applying query optimization techniques specifically adapted for LLMs. We introduce a database system, GALOIS, that sits between the query and the LLM, effectively using the latter as a storage layer. We design alternative physical operators tailored for LLM-based query execution and adapt traditional optimization strategies to this novel context. For example, while pushing down operators in the query plan reduces execution cost (fewer calls to the model), it might complicate the call to the LLM and deteriorate result quality. Additionally, these models lack a traditional catalog for optimization, leading us to develop methods to dynamically gather such metadata during query execution. Our solution is compatible with any LLM and balances the trade-off between query result quality and execution cost. Experiments show up to 144% quality improvement over questions in Natural Language and 29% over direct SQL execution, highlighting the advantages of integrating database solutions with LLMs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Abacus: A Cost-Based Optimizer for Semantic Operator SystemsMatthew Russo, Chunwei Liu, Sivaprasad Sudhir, Gerardo Vitagliano 等VLDB 2026 · 被引用 9 次
- 100x Cost & Latency Reduction: Performance Analysis of AI Query Approximation using Lightweight Proxy Models: [Experiments & Analysis]Yeounoh Chung, Rushabh Desai, Jian He, Yu Xiao 等SIGMOD 2026 · 被引用 8 次
- Relational Deep Dive: Error-Aware Queries Over Unstructured DataDaren Chao, Kaiwen Chen, Naiqing Guan, Nick KoudasVLDB 2026 · 被引用 3 次
- Factorized and Vectorized Execution: Optimizing Analytical and Semantic Queries over RelationsSunny Yasser, Anas Dorbani, Amine MhedhbiSIGMOD 2026 · 被引用 1 次
- Prefix-Cache-Aware Data Reordering for LLM-Augmented Database AnalyticsYingze Li, dong wang, Yiming Guo, Yao Chen 等ICML 2026
它引用的顶会 Paper17
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 被引用 3,228 次
- TURL: Table Understanding through Representation LearningXiang Deng, Huan Sun, Alyssa Lees, You Wu 等VLDB 2021 · 被引用 2,406 次
- Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMsMiao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li 等ICLR 2024 · 被引用 867 次
相关 Paper
- Can Large Language Models Be Query Optimizer for Relational Databases?Jie Tan, Kangfei Zhao, Rui Li, Jeffrey Xu Yu 等SIGMOD 2026 · 被引用 6 次
- Weaver: Interweaving SQL and LLM for Table ReasoningRohit Khoja, Devanshu Gupta, Yanjie Fu, Dan Roth 等EMNLP 2025 · 被引用 1 次
- QUEST: Query Optimization in Unstructured Document AnalysisZhaoze Sun, Chengliang Chai, Qiyan Deng, Kaisen Jin 等VLDB 2025 · 被引用 9 次
- SEMA: A High-performance System for LLM-based Semantic Query ProcessingKangkang Qi, Dongyang Xie, Wenbo Li, Hao Zhang 等VLDB 2026 · 被引用 5 次
- BOND: A Co-Designed Framework for LLM-Powered Analytics Over Relational DataLixiang Chen, Qin Zheng, Zhicheng Pan, Chengcheng Yang 等ICDE 2026
