Automated Program Repair, What Is It Good For? Not Absolutely Nothing!
Hadeel Eladawy, Claire Le Goues, Yuriy Brun
Abstract
Industrial deployments of automated program repair (APR), e.g., at Facebook and Bloomberg, signal a new milestone for this exciting and potentially impactful technology. In these deployments, developers use APR-generated patch suggestions as part of a human-driven debugging process. Unfortunately, little is known about how using patch suggestions affects developers during debugging. This paper conducts a controlled user study with 40 developers with a median of 6 years of experience. The developers engage in debugging tasks on nine naturally-occurring defects in real-world, open-source, Java projects, using Recoder, SimFix, and TBar, three state-of-the-art APR tools. For each debugging task, the developers either have access to the project's tests, or, also, to code suggestions that make all the tests pass. These suggestions are either developer-written or APR-generated, which can be correct or deceptive. Deceptive suggestions, which are a common APR occurrence, make all the available tests pass but fail to generalize to the intended specification. Through a total of 160 debugging sessions, we find that access to a code suggestion significantly increases the odds of submitting a patch. Access to correct APR suggestions increase the odds of debugging success by 14,000% as compared to having access only to tests, but access to deceptive suggestions decrease the odds of success by 65%. Correct suggestions also speed up debugging. Surprisingly, we observe no significant difference in how novice and experienced developers are affected by APR, suggesting that APR may find uses across the experience spectrum. Overall, developers come away with a strong positive impression of APR, suggesting promise for APR-mediated, human-driven debugging, despite existing challenges in APR-generated repair quality.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ba844546-1e7d-4c38-b50e-d9f43f4abb6eCited by top-tier papers8
- Seeing is Fixing: Cross-Modal Reasoning with Multimodal LLMs for Visual Software Issue RepairKai Huang, Jian Zhang, Xiaofei Xie, Chunyang ChenASE 2025 · 5 citations
- Rango: Adaptive Retrieval-Augmented Proving for Automated Software VerificationKyle Thompson, Nuno Saavedra, Pedro Carrott, Kevin Fisher et al.ICSE 2025 · 4 citations
- ChatDBG: Augmenting Debugging with Large Language ModelsKyla Levin, Nicolas van Kempen, Emery D. Berger, Stephen N. FreundFSE 2025 · 3 citations
- QEDCartographer: Automating Formal Verification Using Reward-Free Reinforcement LearningAlex Sanchez-Stern, Abhishek Varghese, Zhanna Kaufman, Shizhuo Dylan Zhang et al.ICSE 2025 · 2 citations
- AdaptEval: A Benchmark for Evaluating Large Language Models on Code Snippet AdaptationTanghaoran Zhang, Xinjun Mao, Shangwen Wang, Yuxin Zhao et al.ASE 2025 · 1 citation
Builds on16
- CoCoNuT: combining context-aware neural translation models using ensemble for program repairThibaud Lutellier, Hung Viet Pham, Lawrence Pang, Yitong Li et al.ISSTA 2020 · 325 citations
- Less training, more repairing please: revisiting automated program repair via zero-shot learningChunqiu Steven Xia, Lingming ZhangFSE 2022 · 223 citations
- A syntax-guided edit decoder for neural program repairQihao Zhu, Zeyu Sun, Yuan-an Xiao, Wenjie Zhang et al.FSE 2021 · 214 citations
- Can automated program repair refine fault localization? a unified debugging approachYiling Lou, Ali Ghanbari, Xia Li, Lingming Zhang et al.ISSTA 2020 · 99 citations
- Baldur: Whole-Proof Generation and Repair with Large Language ModelsEmily First, Markus N. Rabe, Talia Ringer, Yuriy BrunFSE 2023 · 89 citations
Related papers
- Automated Repair of Programs from Large Language ModelsZhiyu Fan, Xiang Gao, Martin Mirchev, Abhik Roychoudhury et al.ICSE 2023 · 213 citations
- Exploring Experiences with Automated Program Repair in PracticeFairuz Nawer Meem, Justin Smith, Brittany JohnsonICSE 2024 · 14 citations
- Evaluating the Impact of Experimental Assumptions in Automated Fault LocalizationEzekiel O. Soremekun, Lukas Kirschner, Marcel Böhme, Mike PapadakisICSE 2023 · 9 citations
- Gamma: Revisiting Template-Based Automated Program Repair Via Mask PredictionQuanjun Zhang, Chunrong Fang, Tongke Zhang, Bowen Yu et al.ASE 2023 · 44 citations
- Better Automatic Program Repair by Using Bug Reports and Tests TogetherManish Motwani, Yuriy BrunICSE 2023 · 22 citations
