Lune

VLDB2026顶会

Multi-Objective Agentic Rewrites for Unstructured Data Processing

Lindsey Linxi Wei, Shreya Shankar, Sepanta Zeighami, Yeounoh Chung, Fatma Özcan, Aditya G. Parameswaran

2026年份
15被引次数

摘要

One year ago, we open-sourced DocETL, a declarative system for LLM-powered data processing that, as of March 2026, has 3.7K GitHub stars and users across journalism, law, medicine, policy, finance, and urban planning. In DocETL, users compose operators described in natural language, with an LLM executing each operator's logic. However, due to complexity in operators or the data, LLMs often give inaccurate results. DocETL addressed this with rewrite directives: abstract rules that guide LLM agents in rewriting pipelines by decomposing operators or data—for example, splitting a single filter("is this email sent from an executive and discussing fraud?") into two separate semantic filters. Yet DocETL optimizes only for accuracy, not cost. How do we optimize for both?

We present MOAR (Multi-Objective Agentic Rewrites), a new optimizer for DocETL. To optimize cost, we add two new categories of directives and extend all three existing ones, bringing the total to over 30—more than doubling DocETL's original set of directives. Moreover, because operators interact unpredictably due to LLM behavior, optimizing them in isolation yields suboptimal plans. So, we design a new global search algorithm that explores rewrites in the context of entire pipelines. Since this space is infinite—every pipeline can be rewritten, and each rewrite rewritten again—we adapt a multi-armed bandit framework to prioritize which pipelines to rewrite. Across six workloads, MOAR achieves 27% higher accuracy than ABACUS, the next-best optimizer, while matching its best accuracy at 55% of its cost.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext 8abcef7b-73d2-4ac7-827a-4a25eaad1242

它引用的顶会 Paper15

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖