Are “Solved Issues” in SWE-bench Really Solved Correctly? An Empirical Study
You Wang, Michael Pradel, Zhongxin Liu
Abstract
Automated issue solving aims to resolve real-world issues in software repositories. The most popular benchmarks for automated issue solving are SWE-bench and its human-filtered subset SWE-bench Verified, which are widely used to evaluate foundation models and software engineering agents. These benchmarks leverage testing to validate generated patches. However, because testing is rarely exhaustive, a patch may pass the tests but nevertheless fail to match the developers’ expectations. Unfortunately, it is currently unclear to what extent evaluations performed with SWE-bench suffer from such plausible but incorrect patches. This paper presents an in-depth empirical study of the correctness of plausible patches generated by three state-of-the-art issue-solving tools (CodeStory, LearnByInteract, and OpenHands) evaluated on SWE-bench Verified. We extensively test and inspect generated patches, and compare them against human-written ground truth patches. The core of our methodology is a novel technique for differential patch testing, called PatchDiff, which automatically exposes behavioral discrepancies between two patches. Our findings reveal critical weaknesses in SWE-bench’s patch validation mechanism, which causes 7.8% of all patches to count as “correct” while failing the developer-written test suite. Moreover, our novel automated technique reveals that even more (29.6%) plausible patches induce different behavior than the ground truth patches. These behavioral differences are often due to similar, but divergent implementations (46.8%) and due to generated patches that adapt more behavior than the ground truth patches (27.3%). Our manual inspection shows that 28.6% of behaviorally divergent patches are certainly incorrect. Combined, the different weaknesses lead to an inflation of reported resolution rates by 6.4 absolute percent points. Our findings are a call to arms for more robust and reliable evaluation of issue-solving tools. We envision our automated differential patch testing technique to be useful for this purpose.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- SPICE: An Automated SWE-Bench Labeling Pipeline for Issue Clarity, Test Coverage, and Effort EstimationGustavo Ansaldi Oliva, Gopi Krishnan Rajbahadur, Aaditya Bhatia, Haoxiang Zhang et al.ASE 2025 · 9 citations
- SWINGARENA: Adversarial Programming Arena for Long-context GitHub Issue SolvingWendong XU, Jing Xiong, Chenyang Zhao, Qiujiang Chen et al.ICLR 2026 · 2 citations
- SWE-Debate: Competitive Multi-Agent Debate for Software Issue ResolutionHan Li, Yuling Shi, Shaoxin Lin, Xiaodong Gu et al.ICSE 2026 · 2 citations
- AutoCodeSherpa: Symbolic Explanations in AI Coding AgentsSungmin Kang, Haifeng Ruan, Abhik RoychoudhuryISSTA 2026
- Names Are All You Need: Effective and Safe Regression Test Selection for PythonYou Wang, Michael Pradel, Zhongxin LiuISSTA 2026
Builds on20
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 2,317 citations
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao et al.ICLR 2024 · 2,082 citations
- SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringJohn Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret et al.NeurIPS 2024 · 2,059 citations
- SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code AgentsNiels Mündler, Mark Niklas Müller, Jingxuan He, Martin T. VechevNeurIPS 2024 · 172 citations
- AutoCodeRover: Autonomous Program ImprovementYuntong Zhang, Haifeng Ruan, Zhiyu Fan, Abhik RoychoudhuryISSTA 2024 · 96 citations
Related papers
- Automated Benchmark Generation for Repository-Level Coding TasksKonstantinos Vergopoulos, Mark Niklas Müller, Martin T. VechevICML 2025
- Otter: Generating Tests from Issues to Validate SWE PatchesToufique Ahmed, Jatin Ganhotra, Rangeet Pan, Avraham Shinnar et al.ICML 2025
- SWE Data Construction, Automatically!Lianghong Guo, Yanlin Wang, Caihua Li, Wei Tao et al.FSE 2026
- Can Agent Fix Agent Issues?Alfin Wijaya Rahardja, Junwei Liu, Weitong Chen, Zhenpeng Chen et al.NeurIPS 2025 · 4 citations
- BenchChecker: Assessing the Credibility of Bug-Fixing Benchmarks for LLMsDi Wu, Xu He, Shu Wang, Kun SunUSENIX Security 2026
