UTBoost: Rigorous Evaluation of Coding Agents on SWE-Bench
Boxi Yu, Yuxuan Zhu, Pinjia He, Daniel Kang
Abstract
The advent of Large Language Models (LLMs) has spurred the development of coding agents for real-world code generation. As a widely used benchmark for evaluating the code generation capabilities of these agents, SWE-Bench uses real-world problems based on GitHub issues and their corresponding pull requests. However, the manually written test cases included in these pull requests are often insufficient, allowing generated patches to pass the tests without resolving the underlying issue. To address this challenge, we introduce UTGenerator, an LLM-driven test case generator that automatically analyzes codebases and dependencies to generate test cases for real-world Python projects. Building on UTGenerator, we propose UTBoost, a comprehensive framework for test case augmentation. In our evaluation, we identified 36 task instances with insufficient test cases and uncovered 345 erroneous patches incorrectly labeled as passed in the original SWE Bench. These corrections, impacting 40.9% of SWE-Bench Lite and 24.4% of SWE-Bench Verified leaderboard entries, yield 18 and 11 ranking changes, respectively.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a0851dae-6f13-4f7d-94ea-48d956a37e13Cited by top-tier papers6
- VeriEquivBench: An Equivalence Score for Ground-Truth-Free Evaluation of Formally Verifiable CodeLingfei Zeng, Fengdi Che, Xuhan Huang, Fei Ye et al.ICLR 2026 · 8 citations
- Pervasive Annotation Errors Break Text-to-SQL Benchmarks and LeaderboardsTengjun Jin, Yoojin Choi, Yuxuan Zhu, Daniel KangVLDB 2026 · 7 citations
- How Much Static Structure Do Code Agents Need? A Study of Deterministic AnchoringZhihao Lin, Mingyi Zhou, Yizhuo Yang, Li LiISSTA 2026
- AgentSuite: Toward More Reliable Agent Evaluation with a Component-Based Benchmark Auditing PipelineHyewon Suh, Binfei Ji, Seojune Lee, Rishi Khare et al.ICML 2026
- SWE-ABS: Adversarial Benchmark Strengthening Exposes Inflated Success Rates on Test-based BenchmarkBoxi Yu, Yang Cao, Yuzhong Zhang, Liting Lin et al.ICML 2026
Builds on4
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 2,317 citations
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao et al.ICLR 2024 · 2,082 citations
- SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringJohn Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret et al.NeurIPS 2024 · 2,059 citations
- RepoGraph: Enhancing AI Software Engineering with Repository-level Code GraphSiru Ouyang, Wenhao Yu, Kaixin Ma, Zilin Xiao et al.ICLR 2025
Related papers
- SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code AgentsNiels Mündler, Mark Niklas Müller, Jingxuan He, Martin T. VechevNeurIPS 2024 · 172 citations
- SWE-Perf: Can Language Models Optimize Code Performance on Real-World Repositories?Xinyi He, Qian Liu, Mingzhe Du, Lin Yan et al.ICML 2026 · 31 citations
- Unified Software Engineering Agent as AI Software EngineerLeonhard Applis, Yuntong Zhang, Shanchao Liang, Nan Jiang et al.ICSE 2026
- Leveraging Large Language Models for Enhancing the Understandability of Generated Unit TestsAmirhossein Deljouyi, Roham Koohestani, Maliheh Izadi, Andy ZaidmanICSE 2025 · 8 citations
- Otter: Generating Tests from Issues to Validate SWE PatchesToufique Ahmed, Jatin Ganhotra, Rangeet Pan, Avraham Shinnar et al.ICML 2025
