SWE-ABS: Adversarial Benchmark Strengthening Exposes Inflated Success Rates on Test-based Benchmark
Boxi Yu, Yang Cao, Yuzhong Zhang, Liting Lin, Junjielong Xu, Zhiqing Zhong, Qinghua Xu, Guancheng Wang, Jialun Cao, Shing-Chi Cheung, Pinjia He, Lionel BRIAND
摘要
The SWE-Bench Verified leaderboard is approaching saturation, with the top system achieving 78.80%. However, we reveal that this performance is inflated: our re-evaluation demonstrates that one in five ``solved'' patches from the top-30 agents are semantically incorrect, passing only because weak test suites fail to expose their errors. We present SWE-ABS, an adversarial framework that strengthens test suites through a two-stage pipeline: (1) coverage-driven augmentation utilizing program slicing to target untested code regions, and (2) mutation-driven adversarial testing that synthesizes plausible-but-incorrect patches to expose semantic blind spots. On SWE-Bench Verified (500 instances), SWE-ABS strengthens 50.2% of instances (a improvement over prior work) and rejects 19.78% of previously passing patches. Consequently, the top agent's score decreases from 78.80% to 62.20%, causing significant leaderboard reshuffling (e.g., the top-ranked agent drops to 5th place).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 被引用 2,317 次
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao 等ICLR 2024 · 被引用 2,082 次
- SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code AgentsNiels Mündler, Mark Niklas Müller, Jingxuan He, Martin T. VechevNeurIPS 2024 · 被引用 172 次
- UTBoost: Rigorous Evaluation of Coding Agents on SWE-BenchBoxi Yu, Yuxuan Zhu, Pinjia He, Daniel KangACL 2025 · 被引用 20 次
- SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?Xiang Deng, Jeff Da, Edwin Pan, Yannis Yiming He 等ICML 2026
相关 Paper
- Are “Solved Issues” in SWE-bench Really Solved Correctly? An Empirical StudyYou Wang, Michael Pradel, Zhongxin LiuICSE 2026 · 被引用 2 次
- Agentic Rubrics as Contextual Verifiers for SWE AgentsMohit Raghavendra, Anisha Gunjal, Bing Liu, Yunzhong HeACL 2026 · 被引用 10 次
- Test vs Mutant: Adversarial LLM Agents for Robust Unit Test GenerationPengyu Chang, Yixiong Fang, Silin Chen, Yuling Shi 等ISSTA 2026
- Automated Benchmark Generation for Repository-Level Coding TasksKonstantinos Vergopoulos, Mark Niklas Müller, Martin T. VechevICML 2025
- Understanding Automated Program Repair Agents through the Lens of Traceability: An Empirical StudyIra Ceka, Hailie Mitchell, Saurabh Pujar, Luca Buratti 等ISSTA 2026
