CentaurEval: Benchmarking Human-in-the-Loop Value in Agentic Coding
Hanjun Luo, Chiming Ni, Jiaheng Wen, Zhimu Huang, Bingduo Liao, Sylvia Chung Yan Shan, Yiran Wang, Yingbin Jin, Jialin Li, Xinfeng Li, Wenyuan Xu, XiaoFeng Wang, Hanan Salam
Abstract
LLM-powered coding agents are reshaping the development paradigm. However, existing evaluation systems, neither traditional tests for humans nor benchmarks for LLMs, fail to capture this shift, excluding problems that require both human reasoning to guide solutions and AI efficiency for implementation. We introduce CentaurEval, a unified, ecologically valid benchmark for measuring human-in-the-loop value in coding. Cen-taurEval's core innovation is its "Collaboration-Necessary" problem templates, which are intractable for standalone LLMs or humans, but solvable through effective collaboration. Centau-rEval dynamically instantiates tasks from 45 templates, providing a standardized IDE for humans and a reproducible 450-task toolkit for LLMs. We benchmark 45 participants against 5 LLMs under 4 levels of human intervention. Results show that while LLMs or humans alone achieve poor pass rates (0.67% and 18.89%), human-AI collaboration significantly improves to 31.11%. Our analysis reveals an emerging co-reasoning partnership, challenging the traditional human-tool hierarchy by showing that strategic breakthroughs can originate from either humans or AI.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on10
- DS-1000: A Natural and Reliable Benchmark for Data Science Code GenerationYuhang Lai, Chengxi Li, Yiming Wang, Tianyi Zhang et al.ICML 2023 · 504 citations
- Grounded Copilot: How Programmers Interact with Code-Generating ModelsShraddha Barke, Michael B. James, Nadia PolikarpovaOOPSLA 2023 · 408 citations
- Using an LLM to Help With Code UnderstandingDaye Nam, Andrew Macvean, Vincent J. Hellendoorn, Bogdan Vasilescu et al.ICSE 2024 · 264 citations
- CoderEval: A Benchmark of Pragmatic Code Generation with Generative Pre-trained ModelsHao Yu, Bo Shen, Dezhi Ran, Jiaxin Zhang et al.ICSE 2024 · 107 citations
- Reading Between the Lines: Modeling User Behavior and Costs in AI-Assisted ProgrammingHussein Mozannar, Gagan Bansal, Adam Fourney, Eric HorvitzCHI 2024 · 88 citations
Related papers
- Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent EvaluationSayash Kapoor, Benedikt Stroebl, Peter Kirgis, Nitya Nadgir et al.ICLR 2026 · 86 citations
- Agent-as-a-Judge: Evaluate Agents with AgentsMingchen Zhuge, Changsheng Zhao, Dylan R. Ashley, Wenyi Wang et al.ICML 2025
- CATArena: Evaluating Evolutionary Capabilities of Code Agents via Iterative TournamentsLingyue Fu, Xin Ding, Linyue Pan, Yaoming Zhu et al.ICML 2026 · 3 citations
- Interactive Evaluation of Large Language Models for Multi-Requirement Software Engineering TasksDimitrios Rontogiannis, Maxime Peyrard, Nicolas Mario Baldwin, Martin Josifoski et al.AAAI 2026 · 1 citation
- Code with Me or for Me? How Increasing AI Automation Transforms Developer WorkflowsValerie Chen, Ameet Talwalkar, Robert Brennan, Graham NeubigCHI 2026 · 2 citations
