VeRO: A Harness for Agents to Optimize Agents
Varun Ursekar, Apaar Shanker, Veronica Chatrath, Yuan Xue, Samuel Denton
摘要
An important emerging application of coding agents is agent harness optimization : the iterative improvement of a target agent by editing and evaluating its code. Despite its relevance, the community lacks a systematic understanding of coding agent performance on this task. Harness optimization differs from conventional software engineering: agent harnesses interleave deterministic code with stochastic LLM completions, requiring structured capture of both intermediate execution traces and downstream outcomes. To address these challenges, we introduce (1) VeRO (Versioning, Rewards, and Observations), an outer harness that provides versioned snapshots, budget-controlled evaluation, and structured execution traces of target harnesses , and (2) VeRO-Bench, a benchmark suite of target agents and tasks with reference evaluation procedures. Using VeRO, we conduct an empirical study comparing optimizers across tasks and analyzing which modifications reliably improve target agent harnesses. We release VeRO to support research on agent optimization as a core capability for coding agents. Code is available at https://github.com/scaleapi/vero.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao 等ICLR 2024 · 被引用 2,082 次
- SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringJohn Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret 等NeurIPS 2024 · 被引用 2,059 次
- Large Language Models as OptimizersChengrun Yang, Xuezhi Wang, Yifeng Lu, Hanxiao Liu 等ICLR 2024 · 被引用 817 次
- GAIA: a benchmark for General AI AssistantsGrégoire Mialon, Clémentine Fourrier, Thomas Wolf, Yann LeCun 等ICLR 2024 · 被引用 716 次
- Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line InterfacesMike A. Merrill, Alexander Glenn Shaw, Nicholas Carlini, Boxuan Li 等ICLR 2026 · 被引用 520 次
相关 Paper
- SecureVibeBench: Benchmarking Secure Vibe Coding of AI Agents via Reconstructing Vulnerability-Introducing ScenariosJunkai Chen, Huihui Huang, Yunbo Lyu, Junwen An 等ACL 2026 · 被引用 5 次
- Unified Software Engineering Agent as AI Software EngineerLeonhard Applis, Yuntong Zhang, Shanchao Liang, Nan Jiang 等ICSE 2026
- OctoBench: Benchmarking Scaffold-Aware Instruction Following in Repository-Grounded Agentic CodingDeming Ding, Shichun Liu, Enhui Yang, Jiahang Lin 等ACL 2026 · 被引用 10 次
- CATArena: Evaluating Evolutionary Capabilities of Code Agents via Iterative TournamentsLingyue Fu, Xin Ding, Linyue Pan, Yaoming Zhu 等ICML 2026 · 被引用 3 次
- From Reproduction to Replication: Evaluating Research Agents with Progressive Code MaskingGyeongwon James Kim, Alex Wilf, Louis-Philippe Morency, Daniel FriedICLR 2026 · 被引用 12 次
