Querying Templatized Document Collections with Large Language Models
Yiming Lin, Madelon Hulsebos, Ruiying Ma, Shreya Shankar, Sepanta Zeighami, Aditya G. Parameswaran, Eugene Wu
摘要
Querying and extracting value from unstructured document collection remains a considerable challenge. While Large Language Models (LLMs) have made remarkable progress in document understanding, they fail to give high accuracy results for analytical queries on documents, and additionally incur high costs. While Retrieval-Augmented Generation (RAG) can reduce costs, accuracy degrades further. Our key insight is that documents in a collection often follow similar templates that impart a common semantic structure. We therefore introduce Zendb, a document analytics system that leverages this semantic structure, coupled with LLMs, to answer ad-hoc SQL queries on document collections. Zendb efficiently extracts semantic hierarchical structures from such templatized documents and introduces a novel query engine that leverages these structures for accurate and cost-effective query execution. Extensive experiments on three real-world document collections demonstrate ZENDB's benefits, achieving up to 31× cost savings compared to LLM-based baselines, while maintaining or improving accuracy, and surpassing RAG-based baselines by up to 61% in precision and 81% in recall, at a marginally higher cost.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper5
- Multi-Objective Agentic Rewrites for Unstructured Data ProcessingLindsey Linxi Wei, Shreya Shankar, Sepanta Zeighami, Yeounoh Chung 等VLDB 2026 · 被引用 15 次
- Task Cascades for Efficient Unstructured Data ProcessingShreya Shankar, Sepanta Zeighami, Aditya G. ParameswaranSIGMOD 2026 · 被引用 7 次
- Relational Deep Dive: Error-Aware Queries Over Unstructured DataDaren Chao, Kaiwen Chen, Naiqing Guan, Nick KoudasVLDB 2026 · 被引用 3 次
- Automated Discovery of Test Oracles for Database Management Systems Using LLMsQiuyang Mang, Runyuan He, Suyang Zhong, Xiaoxuan Liu 等SIGMOD 2026 · 被引用 1 次
- HCT-QA: A Benchmark for Question Answering on Human-Centric TablesMohammad Shahmeer Ahmad, Zan Ahmad Naeem, Michaël Aupetit, Ahmed K. Elmagarmid 等ICDE 2026
相关 Paper
- QUEST: Query Optimization in Unstructured Document AnalysisZhaoze Sun, Chengliang Chai, Qiyan Deng, Kaisen Jin 等VLDB 2025 · 被引用 9 次
- SEMA: A High-performance System for LLM-based Semantic Query ProcessingKangkang Qi, Dongyang Xie, Wenbo Li, Hao Zhang 等VLDB 2026 · 被引用 5 次
- SchemaRAG: A Schema-aware Retrieval-Augmented Generation Framework for Text-to-SQLDi Wu, Zetong Tang, Yi He, Xin LuoSIGMOD 2026 · 被引用 9 次
- Reliable Answers for Recurring Questions: Boosting Text-to-SQL Accuracy with Template Constrained DecodingSmit Jivani, Sarvam Maheshwari, Sunita SarawagiSIGMOD 2026 · 被引用 1 次
- BookRAG: A Hierarchical Structure-aware Index-based Approach for Retrieval-Augmented Generation on Complex DocumentsShu Wang, Yingli Zhou, Yixiang FangVLDB 2026 · 被引用 16 次
