SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Xiang Deng, Jeff Da, Edwin Pan, Yannis Yiming He, Charles Ide, Kanak Garg, Niklas Lauffer, Andrew Park, Chetan Rane, Karmini Sampath, Maya Krishnan, Srivatsa Kundurthy
摘要
We present SWE-Bench Pro, a comprehensive benchmark designed to evaluate software engineering capabilities through complex, realistic programming challenges. This benchmark extends beyond traditional algorithmic problems to encompass the full spectrum of professional software development tasks. The dataset comprises 1,865 problems sourced from 41 active software engineering repositories, spanning 123 unique programming languages and various application domains. The benchmark is structured into public and private components, with public access to problems from 11 repositories and private evaluation sets from 12 repositories across 4 distinct problem categories. SWE-Bench Pro addresses limitations of existing evaluation frameworks by incorporating problems that reflect real-world software engineering scenarios, including substantial codebases, complex enterprise applications, and multi-file projects requiring sophisticated reasoning and code modification skills. Problems range from early-stage startup environments to enterprise-level applications, with the private commercial set remaining inaccessible to maintain evaluation integrity while enabling public access to representative problems for professional development. Our evaluation methodology employs diverse coding approaches and models under controlled conditions, ensuring robust performance assessment across multiple programming paradigms. Results demonstrate significant performance variations across different problem categories, with traditional algorithmic challenges showing notably higher success rates compared to complex, multi-file engineering tasks. The benchmark reveals substantial gaps in current capabilities for handling real-world software engineering scenarios, particularly in areas requiring deep contextual understanding, cross-file reasoning, and integration with existing large-scale systems. This work contributes a more comprehensive and realistic evaluation framework for assessing software engineering capabilities, providing insights into current limitations and establishing a foundation for future development in automated software engineering tools and methodologies.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- Toward Training Superintelligent Software Agents through Self-Play SWE-RLYuxiang Wei, Zhiqing Sun, Emily McMilin, Jonas Gehring 等ICML 2026 · 被引用 32 次
- SWE-rebench V2: Language-Agnostic SWE Task Collection at ScaleIbragim Badertdinov, Maksim Nekrashevich, Anton Shevtsov, Aleksandr GolubevICML 2026 · 被引用 13 次
- EDINET-Bench: Evaluating LLMs on Complex Financial Tasks using Japanese Financial StatementsIssa Sugiura, Takashi Ishida, Taro Makino, Chieko Tazuke 等ICLR 2026 · 被引用 9 次
- ReCode: Updating Code API Knowledge with Reinforcement LearningHaoze Wu, Yunzhi Yao, Wenhao Yu, Ningyu ZhangAAAI 2026 · 被引用 7 次
- Token-Level LLM Collaboration via FusionRouteNuoya Xiong, Yuhang Zhou, Hanqing Zeng, Zhaorun Chen 等ICML 2026 · 被引用 7 次
它引用的顶会 Paper4
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao 等ICLR 2024 · 被引用 2,082 次
- SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringJohn Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret 等NeurIPS 2024 · 被引用 2,059 次
- AutoCodeRover: Autonomous Program ImprovementYuntong Zhang, Haifeng Ruan, Zhiyu Fan, Abhik RoychoudhuryISSTA 2024 · 被引用 96 次
相关 Paper
- SWE-Compass: Towards Unified Evaluation of Agentic Coding Abilities for Large Language ModelsJingxuan Xu, Ken Deng, Weihao Li, Songwei Yu 等ICML 2026 · 被引用 9 次
- SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?John Yang, Carlos E. Jimenez, Alex L. Zhang, Kilian Lieret 等ICLR 2025
- DSBench: How Far Are Data Science Agents from Becoming Data Science Experts?Liqiang Jing, Zhehui Huang, Xiaoyang Wang, Wenlin Yao 等ICLR 2025
- CoreCodeBench: Decoupling Code Intelligence via Fine-Grained Repository-Level TasksLingyue Fu, Hao Guan, Bolun Zhang, Haowei Yuan 等ACL 2026
- CodeMMLU: A Multi-Task Benchmark for Assessing Code Understanding & Reasoning Capabilities of CodeLLMsDung Manh Nguyen, Thang Chau Phan, Nam Le Hai, Tien-Thong Doan 等ICLR 2025
