Multi-Objective Agentic Rewrites for Unstructured Data Processing
Lindsey Linxi Wei, Shreya Shankar, Sepanta Zeighami, Yeounoh Chung, Fatma Özcan, Aditya G. Parameswaran
摘要
One year ago, we open-sourced DocETL, a declarative system for LLM-powered data processing that, as of March 2026, has 3.7K GitHub stars and users across journalism, law, medicine, policy, finance, and urban planning. In DocETL, users compose operators described in natural language, with an LLM executing each operator's logic. However, due to complexity in operators or the data, LLMs often give inaccurate results. DocETL addressed this with rewrite directives: abstract rules that guide LLM agents in rewriting pipelines by decomposing operators or data—for example, splitting a single filter("is this email sent from an executive and discussing fraud?") into two separate semantic filters. Yet DocETL optimizes only for accuracy, not cost. How do we optimize for both?
We present MOAR (Multi-Objective Agentic Rewrites), a new optimizer for DocETL. To optimize cost, we add two new categories of directives and extend all three existing ones, bringing the total to over 30—more than doubling DocETL's original set of directives. Moreover, because operators interact unpredictably due to LLM behavior, optimizing them in isolation yields suboptimal plans. So, we design a new global search algorithm that explores rewrites in the context of entire pipelines. Since this space is infinite—every pipeline can be rewritten, and each rewrite rewritten again—we adapt a multi-armed bandit framework to prioritize which pipelines to rewrite. Across six workloads, MOAR achieves 27% higher accuracy than ABACUS, the next-best optimizer, while matching its best accuracy at 55% of its cost.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- DSPy: Compiling Declarative Language Model Calls into State-of-the-Art PipelinesOmar Khattab, Arnav Singhvi, Paridhi Maheshwari, Zhiyuan Zhang 等ICLR 2024 · 被引用 170 次
- Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data LakesSimran Arora, Brandon Yang, Sabri Eyuboglu, Avanika Narayan 等VLDB 2024 · 被引用 165 次
- DocETL: Agentic Query Rewriting and Evaluation for Complex Document ProcessingShreya Shankar, Tristan Chambers, Tarak Shah, Aditya G. Parameswaran 等VLDB 2025 · 被引用 62 次
相关 Paper
- Abacus: A Cost-Based Optimizer for Semantic Operator SystemsMatthew Russo, Chunwei Liu, Sivaprasad Sudhir, Gerardo Vitagliano 等VLDB 2026 · 被引用 9 次
- Escaping Whack-a-Mole: Optimizing Documentation as Repo-Specific Playbooks for Coding AgentsYutong Cheng, Haifeng Chen, Wenchao Yu, Xujiang Zhao 等ICML 2026
- ViseGPT: Towards Better Alignment of LLM-generated Data Wrangling Scripts and User PromptsJiajun Zhu, Xinyu Cheng, Zhongsu Luo, Yunfan Zhou 等UIST 2025 · 被引用 1 次
- ReSequel: Robust LLM-assisted Query Rewriting and Optimization using Templatization and SamplingSaeed Fathollahzadeh, Essam Mansour, Matthias BoehmVLDB 2026
- Semantic Operators and Their Optimization: Towards AI-Based Data Analytics with Accuracy GuaranteesLiana Patel, Siddharth Jha, Melissa Z. Pan, Harshit Gupta 等VLDB 2025 · 被引用 16 次
