Demystifying LLM-Based Software Engineering Agents
Chunqiu Steven Xia, Yinlin Deng, Soren Dunn, Lingming Zhang
摘要
Recent advancements in large language models (LLMs) have significantly advanced the automation of software development tasks, including code synthesis, program repair, and test generation. More recently, researchers and industry practitioners have developed various autonomous LLM agents to perform end-to-end software development tasks. These agents are equipped with the ability to use tools, run commands, observe feedback from the environment, and plan for future actions. However, the complexity of these agent-based approaches, together with the limited abilities of current LLMs, raises the following question: Do we really have to employ complex autonomous software agents? To attempt to answer this question, we build Agentless -an agentless approach to automatically resolve software development issues. Compared to the verbose and complex setup of agent-based approaches, Agentless employs a simplistic three-phase process of localization, repair, and patch validation, without letting the LLM decide future actions or operate with complex tools. Our results on the popular SWE-bench Lite benchmark show that surprisingly the simplistic Agentless is able to achieve both the highest performance (32.00%, 96 correct fixes) and low cost ($0.70) compared with all existing open-source software agents at the time of paper submission! Agentless also achieves more than 50% solve rate when using Claude 3.5 Sonnet on the new SWE-bench Verified benchmark. In fact, Agentless has already been adopted by OpenAI as the go-to approach to showcase the real-world coding performance of both GPT-4o and the new o1 models; more recently, Agentless has also been used by DeepSeek to evaluate their newest DeepSeek V3 and R1 models. Furthermore, we manually classified the problems in SWE-bench Lite and found problems with exact ground truth patches or insufficient/misleading issue descriptions. As such, we construct SWE-bench Lite-𝑆 by excluding such problematic issues to perform more rigorous evaluation and comparison. Our work highlights the currently overlooked potential of a simplistic, cost-effective technique in autonomous software development. We hope Agentless will help reset the baseline, starting point, and horizon for autonomous software agents, and inspire future work along this crucial direction. We have open-sourced Agentless at: https://github.com/OpenAutoCoder/Agentless CCS Concepts: • Software and its engineering → Software testing and debugging;
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper38
- Toward Training Superintelligent Software Agents through Self-Play SWE-RLYuxiang Wei, Zhiqing Sun, Emily McMilin, Jonas Gehring 等ICML 2026 · 被引用 32 次
- Is Vibe Coding Safe? Benchmarking Vulnerability of Agent-Generated Code in Real-World TasksSongwen Zhao, Danqing Wang, Kexun Zhang, Jiaxuan Luo 等ICML 2026 · 被引用 22 次
- Improving Code Localization with Repository MemoryBoshi Wang, Weijian Xu, Yunsheng Li, Xuemei Gao 等ICLR 2026 · 被引用 20 次
- Huxley-Göodel Machine: Human-Level Coding Agent Development by an Approximation of the Optimal Self-Improving MachineWenyi Wang, Piotr Piekos, Li Nanbo, Firas Laakom 等ICLR 2026 · 被引用 15 次
- EvoClaw: Evaluating AI Agents on Continuous Software EvolutionGangda Deng, Zhaoling Chen, Zhongming Yu, Haoyang Fan 等ICML 2026 · 被引用 6 次
它引用的顶会 Paper29
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao 等ICLR 2024 · 被引用 2,082 次
- SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringJohn Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret 等NeurIPS 2024 · 被引用 2,059 次
- Large Language Models Can Be Easily Distracted by Irrelevant ContextFreda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales 等ICML 2023 · 被引用 970 次
- Large Language Models Cannot Self-Correct Reasoning YetJie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng 等ICLR 2024 · 被引用 858 次
- Automated Program Repair in the Era of Large Pre-trained Language ModelsChunqiu Steven Xia, Yuxiang Wei, Lingming ZhangICSE 2023 · 被引用 321 次
相关 Paper
- Kimi-Dev: Agentless Training as Skill Prior for SWE-agentsZonghan Yang, Shengjie Wang, Kelin Fu, Wenyang He 等ICLR 2026 · 被引用 34 次
- AutoCodeRover: Autonomous Program ImprovementYuntong Zhang, Haifeng Ruan, Zhiyu Fan, Abhik RoychoudhuryISSTA 2024 · 被引用 96 次
- FeatureBench: Benchmarking Agentic Coding for Complex Feature DevelopmentQixing Zhou, Jiacheng Zhang, Haiyang Wang, Rui Hao 等ICLR 2026 · 被引用 30 次
- OrcaLoca: An LLM Agent Framework for Software Issue LocalizationZhongming Yu, Hejia Zhang, Yujie Zhao, Hanxian Huang 等ICML 2025
- SWE-GPT: A Process-Centric Language Model for Automated Software ImprovementYingwei Ma, Rongyu Cao, Yongchang Cao, Yue Zhang 等ISSTA 2025 · 被引用 1 次
