Are “Solved Issues” in SWE-bench Really Solved Correctly? An Empirical Study
You Wang, Michael Pradel, Zhongxin Liu
摘要
Automated issue solving aims to resolve real-world issues in software repositories. The most popular benchmarks for automated issue solving are SWE-bench and its human-filtered subset SWE-bench Verified, which are widely used to evaluate foundation models and software engineering agents. These benchmarks leverage testing to validate generated patches. However, because testing is rarely exhaustive, a patch may pass the tests but nevertheless fail to match the developers’ expectations. Unfortunately, it is currently unclear to what extent evaluations performed with SWE-bench suffer from such plausible but incorrect patches. This paper presents an in-depth empirical study of the correctness of plausible patches generated by three state-of-the-art issue-solving tools (CodeStory, LearnByInteract, and OpenHands) evaluated on SWE-bench Verified. We extensively test and inspect generated patches, and compare them against human-written ground truth patches. The core of our methodology is a novel technique for differential patch testing, called PatchDiff, which automatically exposes behavioral discrepancies between two patches. Our findings reveal critical weaknesses in SWE-bench’s patch validation mechanism, which causes 7.8% of all patches to count as “correct” while failing the developer-written test suite. Moreover, our novel automated technique reveals that even more (29.6%) plausible patches induce different behavior than the ground truth patches. These behavioral differences are often due to similar, but divergent implementations (46.8%) and due to generated patches that adapt more behavior than the ground truth patches (27.3%). Our manual inspection shows that 28.6% of behaviorally divergent patches are certainly incorrect. Combined, the different weaknesses lead to an inflation of reported resolution rates by 6.4 absolute percent points. Our findings are a call to arms for more robust and reliable evaluation of issue-solving tools. We envision our automated differential patch testing technique to be useful for this purpose.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- SPICE: An Automated SWE-Bench Labeling Pipeline for Issue Clarity, Test Coverage, and Effort EstimationGustavo Ansaldi Oliva, Gopi Krishnan Rajbahadur, Aaditya Bhatia, Haoxiang Zhang 等ASE 2025 · 被引用 9 次
- SWINGARENA: Adversarial Programming Arena for Long-context GitHub Issue SolvingWendong XU, Jing Xiong, Chenyang Zhao, Qiujiang Chen 等ICLR 2026 · 被引用 2 次
- SWE-Debate: Competitive Multi-Agent Debate for Software Issue ResolutionHan Li, Yuling Shi, Shaoxin Lin, Xiaodong Gu 等ICSE 2026 · 被引用 2 次
- AutoCodeSherpa: Symbolic Explanations in AI Coding AgentsSungmin Kang, Haifeng Ruan, Abhik RoychoudhuryISSTA 2026
- Names Are All You Need: Effective and Safe Regression Test Selection for PythonYou Wang, Michael Pradel, Zhongxin LiuISSTA 2026
它引用的顶会 Paper20
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 被引用 2,317 次
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao 等ICLR 2024 · 被引用 2,082 次
- SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringJohn Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret 等NeurIPS 2024 · 被引用 2,059 次
- SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code AgentsNiels Mündler, Mark Niklas Müller, Jingxuan He, Martin T. VechevNeurIPS 2024 · 被引用 172 次
- AutoCodeRover: Autonomous Program ImprovementYuntong Zhang, Haifeng Ruan, Zhiyu Fan, Abhik RoychoudhuryISSTA 2024 · 被引用 96 次
相关 Paper
- Automated Benchmark Generation for Repository-Level Coding TasksKonstantinos Vergopoulos, Mark Niklas Müller, Martin T. VechevICML 2025
- Otter: Generating Tests from Issues to Validate SWE PatchesToufique Ahmed, Jatin Ganhotra, Rangeet Pan, Avraham Shinnar 等ICML 2025
- SWE Data Construction, Automatically!Lianghong Guo, Yanlin Wang, Caihua Li, Wei Tao 等FSE 2026
- Can Agent Fix Agent Issues?Alfin Wijaya Rahardja, Junwei Liu, Weitong Chen, Zhenpeng Chen 等NeurIPS 2025 · 被引用 4 次
- BenchChecker: Assessing the Credibility of Bug-Fixing Benchmarks for LLMsDi Wu, Xu He, Shu Wang, Kun SunUSENIX Security 2026
