LLM-based Agents for Automated Bug Fixing: How Far Are We?
Xiangxin Meng, Zexiong Ma, Pengfei Gao, Chao Peng
Abstract
Large language models (LLMs) and LLM-based Agents have been applied to fix bugs automatically, demonstrating the capability in addressing software defects by engaging in development environment interaction, iterative validation and code modification. However, systematic analysis of these agent systems remain limited, particularly regarding performance variations among top-performing ones. In this paper, we examine six repair systems on the SWEbench Verified benchmark for automated bug fixing. We first assess each system's overall performance, noting the instances solvable by all or none of these systems, and explore the capabilities of different systems. We also compare fault localization accuracy at file and code symbol levels and evaluate bug reproduction capabilities. Through analysis, we concluded that further optimization is needed in both the LLM capability itself and the design of Agentic flow to improve the effectiveness of the Agent in bug fixing.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e82d536e-efa2-4ada-bc74-5c8536edcc0dCited by top-tier papers1
Ask how each one uses itBuilds on25
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 2,317 citations
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao et al.ICLR 2024 · 2,082 citations
- SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringJohn Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret et al.NeurIPS 2024 · 2,059 citations
- Automated Program Repair in the Era of Large Pre-trained Language ModelsChunqiu Steven Xia, Yuxiang Wei, Lingming ZhangICSE 2023 · 321 citations
- Magicoder: Empowering Code Generation with OSS-InstructYuxiang Wei, Zhe Wang, Jiawei Liu, Yifeng Ding et al.ICML 2024 · 246 citations
Related papers
- Understanding Automated Program Repair Agents through the Lens of Traceability: An Empirical StudyIra Ceka, Hailie Mitchell, Saurabh Pujar, Luca Buratti et al.ISSTA 2026
- PATCHAGENT: A Practical Program Repair Agent Mimicking Human ExpertiseZheng Yu, Ziyi Guo, Yuhang Wu, Jiahao Yu et al.USENIX Security 2025
- SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code AgentsNiels Mündler, Mark Niklas Müller, Jingxuan He, Martin T. VechevNeurIPS 2024 · 172 citations
- RepairAgent: An Autonomous, LLM-Based Agent for Program RepairIslem Bouzenia, Premkumar T. Devanbu, Michael PradelICSE 2025 · 54 citations
- Taming System Complexity: Demystifying Software Engineering Agents in Diagnosing Linux Kernel FaultsZhenhao Zhou, Zhuochen Huang, Yike He, Chong Wang et al.ACL 2026 · 5 citations
