Understanding Automated Program Repair Agents through the Lens of Traceability: An Empirical Study
Ira Ceka, Hailie Mitchell, Saurabh Pujar, Luca Buratti, Shyam Ramji, Junfeng Yang, Gail Kaiser, Baishakhi Ray
Abstract
Automated Program Repair (APR) agents leverage large language models (LLMs) to autonomously diagnose and patch software bugs using planning, reasoning, and tools. Although these agents show strong performance on leaderboards such as SWE-bench, little is understood about how they take actions, where they fail, and how their behavior compares to human developers. In this paper, we present the first systematic analysis of these limitations using 5 state-of-the-art APR agents. We trace the full decision-making pipelines of the 5 APR agents across 500 real-world repair tasks, from issue description to patch validation.
Our study reveals that, while agents excel at simple fixes, they struggle with logic-intensive bugs, often generating verbose, overfitted patches that pass existing test suites without solving the root cause. Test generation and regression test selection remain major bottlenecks, as agents fail to reproduce issues or run relevant regression tests. Moreover, many agents operate with primitive tooling (e.g. bash scripts) and do not have access to debuggers or program analysis tools. These findings highlight key limitations of current APR systems and motivate several directions for next-generation APR design, including but not limited to:
(1) a shift-left approach emphasizing early, high-quality test generation and validation to reduce spurious fixes and improve semantic correctness; (2) richer, more integrated tool ecosystems; (3) diversified agent architectures that combine complementary strengths; and (4) benchmarks that prioritize semantic repair quality and test-generation fidelity over surface-level success metrics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f472e1d1-d1c0-4165-b789-3cba18ab7d43Cited by top-tier papers2
- Process-Centric Analysis of Agentic Software SystemsShuyang Liu, Yang Chen, Rahul Krishna, Saurabh Sinha et al.OOPSLA 2026 · 1 citation
- BenchChecker: Assessing the Credibility of Bug-Fixing Benchmarks for LLMsDi Wu, Xu He, Shu Wang, Kun SunUSENIX Security 2026
Builds on21
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao et al.ICLR 2024 · 2,082 citations
- SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringJohn Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret et al.NeurIPS 2024 · 2,059 citations
- Grounded Copilot: How Programmers Interact with Code-Generating ModelsShraddha Barke, Michael B. James, Nadia PolikarpovaOOPSLA 2023 · 408 citations
- SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software EvolutionYuxiang Wei, Olivier Duchenne, Jade Copet, Quentin Carbonneaux et al.NeurIPS 2025 · 291 citations
- SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code AgentsNiels Mündler, Mark Niklas Müller, Jingxuan He, Martin T. VechevNeurIPS 2024 · 172 citations
Related papers
- LLM-based Agents for Automated Bug Fixing: How Far Are We?Xiangxin Meng, Zexiong Ma, Pengfei Gao, Chao PengICSE 2026
- PATCHAGENT: A Practical Program Repair Agent Mimicking Human ExpertiseZheng Yu, Ziyi Guo, Yuhang Wu, Jiahao Yu et al.USENIX Security 2025
- Synthetic Repo-level Bug Dataset for Training Automated Program Repair ModelsMinh V. T. Pham, Huy N. Phan, Nhat Hoang Phan, Cuong Chi Le et al.ICSE 2026
- Copiloting the Copilots: Fusing Large Language Models with Completion Engines for Automated Program RepairYuxiang Wei, Chunqiu Steven Xia, Lingming ZhangFSE 2023 · 111 citations
- Automated Program Repair in the Era of Large Pre-trained Language ModelsChunqiu Steven Xia, Yuxiang Wei, Lingming ZhangICSE 2023 · 321 citations
