SWE-bench: Can Language Models Resolve Real-world Github Issues?
Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, Karthik R. Narasimhan
摘要
Language models have outpaced our ability to evaluate them effectively, but for their future development it is essential to study the frontier of their capabilities. We find real-world software engineering to be a rich, sustainable, and challenging testbed for evaluating the next generation of language models. To this end, we introduce SWE-bench, an evaluation framework consisting of software engineering problems drawn from real GitHub issues and corresponding pull requests across popular Python repositories. Given a codebase along with a description of an issue to be resolved, a language model is tasked with editing the codebase to address the issue. Resolving issues in SWE-bench frequently requires understanding and coordinating changes across multiple functions, classes, and even files simultaneously, calling for models to interact with execution environments, process extremely long contexts and perform complex reasoning that goes far beyond traditional code generation tasks. Our evaluations show that both state-of-the-art proprietary models and our fine-tuned model SWE-Llama can resolve only the simplest issues. The best-performing model, Claude 2, is able to solve a mere % of the issues. Advances on SWE-bench represent steps towards LMs that are more practical, intelligent, and autonomous.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper490
- SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringJohn Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret 等NeurIPS 2024 · 被引用 2,059 次
- Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line InterfacesMike A. Merrill, Alexander Glenn Shaw, Nicholas Carlini, Boxuan Li 等ICLR 2026 · 被引用 520 次
- MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse AttentionHuiqiang Jiang, Yucheng Li, Chengruidong Zhang, Qianhui Wu 等NeurIPS 2024 · 被引用 479 次
- SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software EvolutionYuxiang Wei, Olivier Duchenne, Jade Copet, Quentin Carbonneaux 等NeurIPS 2025 · 被引用 291 次
- AREAL: A Large-Scale Asynchronous Reinforcement Learning System for Language ReasoningWei Fu, Jiaxuan Gao, Xujie Shen, Chen Zhu 等NeurIPS 2025 · 被引用 273 次
它引用的顶会 Paper16
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 被引用 2,317 次
- WebArena: A Realistic Web Environment for Building Autonomous AgentsShuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou 等ICLR 2024 · 被引用 1,197 次
- AgentBench: Evaluating LLMs as AgentsXiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu 等ICLR 2024 · 被引用 748 次
- Less training, more repairing please: revisiting automated program repair via zero-shot learningChunqiu Steven Xia, Lingming ZhangFSE 2022 · 被引用 223 次
相关 Paper
- SWE-Compass: Towards Unified Evaluation of Agentic Coding Abilities for Large Language ModelsJingxuan Xu, Ken Deng, Weihao Li, Songwei Yu 等ICML 2026 · 被引用 9 次
- SWE-Perf: Can Language Models Optimize Code Performance on Real-World Repositories?Xinyi He, Qian Liu, Mingzhe Du, Lin Yan 等ICML 2026 · 被引用 31 次
- Training Software Engineering Agents and Verifiers with SWE-GymJiayi Pan, Xingyao Wang, Graham Neubig, Navdeep Jaitly 等ICML 2025
- SWE-GPT: A Process-Centric Language Model for Automated Software ImprovementYingwei Ma, Rongyu Cao, Yongchang Cao, Yue Zhang 等ISSTA 2025 · 被引用 1 次
- SWE Data Construction, Automatically!Lianghong Guo, Yanlin Wang, Caihua Li, Wei Tao 等FSE 2026
