SWE-ABS: Adversarial Benchmark Strengthening Exposes Inflated Success Rates on Test-based Benchmark
Boxi Yu, Yang Cao, Yuzhong Zhang, Liting Lin, Junjielong Xu, Zhiqing Zhong, Qinghua Xu, Guancheng Wang, Jialun Cao, Shing-Chi Cheung, Pinjia He, Lionel BRIAND
Abstract
The SWE-Bench Verified leaderboard is approaching saturation, with the top system achieving 78.80%. However, we reveal that this performance is inflated: our re-evaluation demonstrates that one in five ``solved'' patches from the top-30 agents are semantically incorrect, passing only because weak test suites fail to expose their errors. We present SWE-ABS, an adversarial framework that strengthens test suites through a two-stage pipeline: (1) coverage-driven augmentation utilizing program slicing to target untested code regions, and (2) mutation-driven adversarial testing that synthesizes plausible-but-incorrect patches to expose semantic blind spots. On SWE-Bench Verified (500 instances), SWE-ABS strengthens 50.2% of instances (a improvement over prior work) and rejects 19.78% of previously passing patches. Consequently, the top agent's score decreases from 78.80% to 62.20%, causing significant leaderboard reshuffling (e.g., the top-ranked agent drops to 5th place).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on6
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 2,317 citations
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao et al.ICLR 2024 · 2,082 citations
- SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code AgentsNiels Mündler, Mark Niklas Müller, Jingxuan He, Martin T. VechevNeurIPS 2024 · 172 citations
- UTBoost: Rigorous Evaluation of Coding Agents on SWE-BenchBoxi Yu, Yuxuan Zhu, Pinjia He, Daniel KangACL 2025 · 20 citations
- SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?Xiang Deng, Jeff Da, Edwin Pan, Yannis Yiming He et al.ICML 2026
Related papers
- Are “Solved Issues” in SWE-bench Really Solved Correctly? An Empirical StudyYou Wang, Michael Pradel, Zhongxin LiuICSE 2026 · 2 citations
- Agentic Rubrics as Contextual Verifiers for SWE AgentsMohit Raghavendra, Anisha Gunjal, Bing Liu, Yunzhong HeACL 2026 · 10 citations
- Test vs Mutant: Adversarial LLM Agents for Robust Unit Test GenerationPengyu Chang, Yixiong Fang, Silin Chen, Yuling Shi et al.ISSTA 2026
- Automated Benchmark Generation for Repository-Level Coding TasksKonstantinos Vergopoulos, Mark Niklas Müller, Martin T. VechevICML 2025
- Understanding Automated Program Repair Agents through the Lens of Traceability: An Empirical StudyIra Ceka, Hailie Mitchell, Saurabh Pujar, Luca Buratti et al.ISSTA 2026
