QUEST: Query Optimization in Unstructured Document Analysis
Zhaoze Sun, Chengliang Chai, Qiyan Deng, Kaisen Jin, Xinyu Guo, Han Han, Ye Yuan, Guoren Wang, Lei Cao
Abstract
Most recently, researchers have started building large language models (LLMs) powered data systems that allow users to analyze unstructured text documents like working with a database because LLMs are very effective in extracting attributes from documents. In such systems, LLM-based extraction operations constitute the performance bottleneck of query execution due to the high monetary cost and slow LLM inference. Existing systems typically borrow the query optimization principles popular in relational databases to produce query execution plans, which unfortunately are ineffective in minimizing LLM cost. To fill this gap, we propose QUEST, which features a bunch of novel optimization strategies for unstructured document analysis. First, we introduce an index-based strategy to minimize the cost of each extraction operation. With this index, QUEST quickly retrieves the text segments relevant to the target attributes and only feeds them to LLMs. Furthermore, we design an evidence-augmented retrieval strategy to reduce the possibility of missing relevant segments. Moreover, we develop an instance-optimized query execution strategy: because the attribute extraction cost could vary significantly document by document, QUEST produces different plans for different documents. For each document, QUEST produces a plan to minimize the frequency of attribute extraction. The innovations include LLM cost-aware operator ordering strategies and an optimized join execution approach that transforms joins into filters. Extensive experiments on 3 real-world datasets demonstrate the superiority of QUEST, achieving 30%-6× cost savings while improving the F1 score by 10% -27% compared with state-of-the-art baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bd0f928e-a19e-4eab-b01f-e655ffaffc3cCited by top-tier papers7
- Multi-Objective Agentic Rewrites for Unstructured Data ProcessingLindsey Linxi Wei, Shreya Shankar, Sepanta Zeighami, Yeounoh Chung et al.VLDB 2026 · 15 citations
- Featurized-Decomposition Join: Low-Cost Semantic Joins with GuaranteesSepanta Zeighami, Shreya Shankar, Aditya G. ParameswaranVLDB 2026 · 11 citations
- AgenticScholar: Agentic Data Management with Pipeline Orchestration for Scholarly CorporaHai Lan, Tingting Wang, Zhifeng Bao, Guoliang Li et al.SIGMOD 2026 · 4 citations
- Unstructured Data Analysis using LLMs: A Comprehensive BenchmarkQiyan Deng, Jianhui Li, Chengliang Chai, Ye Yuan et al.VLDB 2026 · 4 citations
- MoDora: Tree-Based Semi-Structured Document Analysis SystemBangrui Xu, Qihang Yao, Zirui Tang, Xuanhe Zhou et al.SIGMOD 2026
Builds on16
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
- Text-to-SQL Empowered by Large Language Models: A Benchmark EvaluationDawei Gao, Haibin Wang, Yaliang Li, Xiuyu Sun et al.VLDB 2024 · 609 citations
- RAPTOR: Recursive Abstractive Processing for Tree-Organized RetrievalParth Sarthi, Salman Abdullah, Aditi Tuli, Shubh Khanna et al.ICLR 2024 · 460 citations
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding SharingPengcheng He, Jianfeng Gao, Weizhu ChenICLR 2023 · 394 citations
Related papers
- Doctopus: Budget-aware Structural Table Extraction from Unstructured DocumentsChengliang Chai, Jiajun Li, Yuhao Deng, Yuanhao Zhong et al.VLDB 2025 · 5 citations
- Can Large Language Models Be Query Optimizer for Relational Databases?Jie Tan, Kangfei Zhao, Rui Li, Jeffrey Xu Yu et al.SIGMOD 2026 · 6 citations
- Logical and Physical Optimizations for SQL Query Execution over Large Language ModelsDario Satriani, Enzo Veltri, Donatello Santoro, Sara Rosato et al.SIGMOD 2025 · 7 citations
- Querying Templatized Document Collections with Large Language ModelsYiming Lin, Madelon Hulsebos, Ruiying Ma, Shreya Shankar et al.ICDE 2025 · 4 citations
- SEMA: A High-performance System for LLM-based Semantic Query ProcessingKangkang Qi, Dongyang Xie, Wenbo Li, Hao Zhang et al.VLDB 2026 · 5 citations
