UTBoost: Rigorous Evaluation of Coding Agents on SWE-Bench
Boxi Yu, Yuxuan Zhu, Pinjia He, Daniel Kang
摘要
The advent of Large Language Models (LLMs) has spurred the development of coding agents for real-world code generation. As a widely used benchmark for evaluating the code generation capabilities of these agents, SWE-Bench uses real-world problems based on GitHub issues and their corresponding pull requests. However, the manually written test cases included in these pull requests are often insufficient, allowing generated patches to pass the tests without resolving the underlying issue. To address this challenge, we introduce UTGenerator, an LLM-driven test case generator that automatically analyzes codebases and dependencies to generate test cases for real-world Python projects. Building on UTGenerator, we propose UTBoost, a comprehensive framework for test case augmentation. In our evaluation, we identified 36 task instances with insufficient test cases and uncovered 345 erroneous patches incorrectly labeled as passed in the original SWE Bench. These corrections, impacting 40.9% of SWE-Bench Lite and 24.4% of SWE-Bench Verified leaderboard entries, yield 18 and 11 ranking changes, respectively.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- VeriEquivBench: An Equivalence Score for Ground-Truth-Free Evaluation of Formally Verifiable CodeLingfei Zeng, Fengdi Che, Xuhan Huang, Fei Ye 等ICLR 2026 · 被引用 8 次
- Pervasive Annotation Errors Break Text-to-SQL Benchmarks and LeaderboardsTengjun Jin, Yoojin Choi, Yuxuan Zhu, Daniel KangVLDB 2026 · 被引用 7 次
- How Much Static Structure Do Code Agents Need? A Study of Deterministic AnchoringZhihao Lin, Mingyi Zhou, Yizhuo Yang, Li LiISSTA 2026
- AgentSuite: Toward More Reliable Agent Evaluation with a Component-Based Benchmark Auditing PipelineHyewon Suh, Binfei Ji, Seojune Lee, Rishi Khare 等ICML 2026
- SWE-ABS: Adversarial Benchmark Strengthening Exposes Inflated Success Rates on Test-based BenchmarkBoxi Yu, Yang Cao, Yuzhong Zhang, Liting Lin 等ICML 2026
它引用的顶会 Paper4
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 被引用 2,317 次
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao 等ICLR 2024 · 被引用 2,082 次
- SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringJohn Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret 等NeurIPS 2024 · 被引用 2,059 次
- RepoGraph: Enhancing AI Software Engineering with Repository-level Code GraphSiru Ouyang, Wenhao Yu, Kaixin Ma, Zilin Xiao 等ICLR 2025
相关 Paper
- SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code AgentsNiels Mündler, Mark Niklas Müller, Jingxuan He, Martin T. VechevNeurIPS 2024 · 被引用 172 次
- SWE-Perf: Can Language Models Optimize Code Performance on Real-World Repositories?Xinyi He, Qian Liu, Mingzhe Du, Lin Yan 等ICML 2026 · 被引用 31 次
- Unified Software Engineering Agent as AI Software EngineerLeonhard Applis, Yuntong Zhang, Shanchao Liang, Nan Jiang 等ICSE 2026
- Leveraging Large Language Models for Enhancing the Understandability of Generated Unit TestsAmirhossein Deljouyi, Roham Koohestani, Maliheh Izadi, Andy ZaidmanICSE 2025 · 被引用 8 次
- Otter: Generating Tests from Issues to Validate SWE PatchesToufique Ahmed, Jatin Ganhotra, Rangeet Pan, Avraham Shinnar 等ICML 2025
