Multi-Objective Agentic Rewrites for Unstructured Data Processing
Lindsey Linxi Wei, Shreya Shankar, Sepanta Zeighami, Yeounoh Chung, Fatma Özcan, Aditya G. Parameswaran
Abstract
One year ago, we open-sourced DocETL, a declarative system for LLM-powered data processing that, as of March 2026, has 3.7K GitHub stars and users across journalism, law, medicine, policy, finance, and urban planning. In DocETL, users compose operators described in natural language, with an LLM executing each operator's logic. However, due to complexity in operators or the data, LLMs often give inaccurate results. DocETL addressed this with rewrite directives: abstract rules that guide LLM agents in rewriting pipelines by decomposing operators or data—for example, splitting a single filter("is this email sent from an executive and discussing fraud?") into two separate semantic filters. Yet DocETL optimizes only for accuracy, not cost. How do we optimize for both?
We present MOAR (Multi-Objective Agentic Rewrites), a new optimizer for DocETL. To optimize cost, we add two new categories of directives and extend all three existing ones, bringing the total to over 30—more than doubling DocETL's original set of directives. Moreover, because operators interact unpredictably due to LLM behavior, optimizing them in isolation yields suboptimal plans. So, we design a new global search algorithm that explores rewrites in the context of entire pipelines. Since this space is infinite—every pipeline can be rewritten, and each rewrite rewritten again—we adapt a multi-armed bandit framework to prioritize which pipelines to rewrite. Across six workloads, MOAR achieves 27% higher accuracy than ABACUS, the next-best optimizer, while matching its best accuracy at 55% of its cost.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8abcef7b-73d2-4ac7-827a-4a25eaad1242Builds on15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- DSPy: Compiling Declarative Language Model Calls into State-of-the-Art PipelinesOmar Khattab, Arnav Singhvi, Paridhi Maheshwari, Zhiyuan Zhang et al.ICLR 2024 · 170 citations
- Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data LakesSimran Arora, Brandon Yang, Sabri Eyuboglu, Avanika Narayan et al.VLDB 2024 · 165 citations
- DocETL: Agentic Query Rewriting and Evaluation for Complex Document ProcessingShreya Shankar, Tristan Chambers, Tarak Shah, Aditya G. Parameswaran et al.VLDB 2025 · 62 citations
Related papers
- Abacus: A Cost-Based Optimizer for Semantic Operator SystemsMatthew Russo, Chunwei Liu, Sivaprasad Sudhir, Gerardo Vitagliano et al.VLDB 2026 · 9 citations
- Escaping Whack-a-Mole: Optimizing Documentation as Repo-Specific Playbooks for Coding AgentsYutong Cheng, Haifeng Chen, Wenchao Yu, Xujiang Zhao et al.ICML 2026
- ViseGPT: Towards Better Alignment of LLM-generated Data Wrangling Scripts and User PromptsJiajun Zhu, Xinyu Cheng, Zhongsu Luo, Yunfan Zhou et al.UIST 2025 · 1 citation
- ReSequel: Robust LLM-assisted Query Rewriting and Optimization using Templatization and SamplingSaeed Fathollahzadeh, Essam Mansour, Matthias BoehmVLDB 2026
- Semantic Operators and Their Optimization: Towards AI-Based Data Analytics with Accuracy GuaranteesLiana Patel, Siddharth Jha, Melissa Z. Pan, Harshit Gupta et al.VLDB 2025 · 16 citations
