DocETL: Agentic Query Rewriting and Evaluation for Complex Document Processing
Shreya Shankar, Tristan Chambers, Tarak Shah, Aditya G. Parameswaran, Eugene Wu
摘要
Analyzing unstructured data has been a persistent challenge in data processing. Large Language Models (LLMs) have shown promise in this regard, leading to recent proposals for declarative frameworks for LLM-powered processing of unstructured data. However, these frameworks focus on reducing cost when executing user-specified operations using LLMs, rather than improving accuracy, executing most operations as-is (in a single LLM call). This is problematic for complex tasks and data, where LLM outputs for user-defined operations are often inaccurate, even with optimized prompts. For example, an LLM may struggle to identify all instances of specific clauses, like force majeure or indemnification, in lengthy legal documents, requiring decomposition of the data, the task, or both. We present DocETL, a system that optimizes complex document processing pipelines, while accounting for LLM shortcomings. Do-cETL offers a declarative interface for users to define such pipelines and uses an agent-based approach to automatically optimize them, leveraging novel agent-based rewrites (that we call rewrite directives), as well as an optimization and evaluation framework. We introduce (i) logical rewriting of pipelines, tailored for LLM-based tasks, (ii) an agent-guided plan evaluation mechanism that synthesizes and orchestrates task-specific validation prompts, and (iii) an optimization algorithm that efficiently finds promising plans, considering the latencies of agent-based plan generation and evaluation. Our evaluation on four different unstructured document analysis tasks demonstrates that DocETL finds plans with outputs that are 21 to 80% more accurate than well-engineered baselines. DocETL is open-source at docetl.org, and as of March 2025, has amassed over 1.7k GitHub Stars, with users spanning a variety of domains.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper28
- Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent EvaluationSayash Kapoor, Benedikt Stroebl, Peter Kirgis, Nitya Nadgir 等ICLR 2026 · 被引用 86 次
- KRAMABENCH: A Benchmark for AI Systems on Data-to-Insight Pipelines over Data LakesEugenie Lai, Gerardo Vitagliano, Ziyu Zhang, Om Chabra 等ICLR 2026 · 被引用 37 次
- SemBench: A Benchmark for Semantic Query Processing EnginesJiale Lao, Andreas Zimmerer, Olga Ovcharenko, Tianji Cong 等VLDB 2026 · 被引用 31 次
- Semantic Operators and Their Optimization: Towards AI-Based Data Analytics with Accuracy GuaranteesLiana Patel, Siddharth Jha, Melissa Z. Pan, Harshit Gupta 等VLDB 2025 · 被引用 16 次
- BookRAG: A Hierarchical Structure-aware Index-based Approach for Retrieval-Augmented Generation on Complex DocumentsShu Wang, Yingli Zhou, Yixiang FangVLDB 2026 · 被引用 16 次
它引用的顶会 Paper16
- Large Language Models Can Be Easily Distracted by Irrelevant ContextFreda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales 等ICML 2023 · 被引用 970 次
- Hard Prompts Made Easy: Gradient-Based Discrete Optimization for Prompt Tuning and DiscoveryYuxin Wen, Neel Jain, John Kirchenbauer, Micah Goldblum 等NeurIPS 2023 · 被引用 454 次
- DSPy: Compiling Declarative Language Model Calls into State-of-the-Art PipelinesOmar Khattab, Arnav Singhvi, Paridhi Maheshwari, Zhiyuan Zhang 等ICLR 2024 · 被引用 170 次
- Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data LakesSimran Arora, Brandon Yang, Sabri Eyuboglu, Avanika Narayan 等VLDB 2024 · 被引用 165 次
- How Large Language Models Will Disrupt Data ManagementRaul Castro Fernandez, Aaron J. Elmore, Michael J. Franklin, Sanjay Krishnan 等VLDB 2023 · 被引用 127 次
相关 Paper
- Multi-Objective Agentic Rewrites for Unstructured Data ProcessingLindsey Linxi Wei, Shreya Shankar, Sepanta Zeighami, Yeounoh Chung 等VLDB 2026 · 被引用 15 次
- Escaping Whack-a-Mole: Optimizing Documentation as Repo-Specific Playbooks for Coding AgentsYutong Cheng, Haifeng Chen, Wenchao Yu, Xujiang Zhao 等ICML 2026
- ViseGPT: Towards Better Alignment of LLM-generated Data Wrangling Scripts and User PromptsJiajun Zhu, Xinyu Cheng, Zhongsu Luo, Yunfan Zhou 等UIST 2025 · 被引用 1 次
- AgentODRL: A Large Language Model-based Multi-agent System for ODRL GenerationWanle Zhong, Keman Huang, Xiaoyong DuAAAI 2026
- DocThinker: Explainable Multimodal Large Language Models with Rule-Based Reinforcement Learning for Document UnderstandingWenwen Yu, Zhibo Yang, Yuliang Liu, Xiang BaiICCV 2025 · 被引用 2 次
