Benchmarking Automated Program Repair: An Extensive Study on Both Real-World and Artificial Bugs
Yicheng Ouyang, Jun Yang, Lingming Zhang
摘要
As bugs are inevitable and prevalent in real-world programs, many Automated Program Repair (APR) techniques have been proposed to generate patches for them. However, due to the lack of a standard for evaluating APR techniques, prior works tend to use different settings and benchmarks in evaluation, threatening the trustworthiness of the evaluation results. Additionally, they typically only adopt plausibility and genuineness as evaluation metrics, which may potentially mask some underlying issues in APR techniques. To overcome these issues, in this paper, we conduct an extensive and multi-dimensional evaluation of nine learning-based and three traditional state-of-the-art APR techniques under the same environment and settings. We employ the widely studied Defects4J V2.0.0 benchmark and a newly constructed large-scale mutation-based benchmark named MuBench, derived from Defects4J and including 1,700 artificial bugs generated by various mutators, to uncover potential limitations in these APR techniques. We also apply multi-dimensional metrics, including compilability/plausibility/genuineness metrics, as well as SYE (SYntactic Equivalence) and TCE (Trivial Compiler Equivalence) metrics, to thoroughly analyze the 1,814,652 generated patches. This paper presents noteworthy findings from the extensive evaluation: Firstly, Large Language Model (LLM) based APR demonstrates less susceptibility to overfitting on the Defects4J V1.2.0 dataset and fixes the most number of bugs. Secondly, the study suggests a promising future for combining traditional and learning-based APR techniques, as they exhibit complementary advantages in fixing different types of bugs. Additionally, this work highlights the necessity for further enhancing patch compilability of learning-based APR techniques, despite the presence of various existing strategies attempting to improve it. The study also reveals other guidelines for enhancing APR techniques, including the need for handling unresolvable symbol compilability issues and reducing duplicate/no-op patch generation. Finally, our study uncovers seven implementation issues in the studied techniques, with five of them confirmed and fixed by the corresponding authors.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- ClassEval-T: Evaluating Large Language Models in Class-Level Code TranslationPengyu Xue, Linhao Wu, Zhen Yang, Chengyi Wang 等ISSTA 2025 · 被引用 5 次
- Enhancing APR with PRISM: A Semantic-Based Approach to Overfitting Patch DetectionDowon Song, Hakjoo OhOOPSLA 2025 · 被引用 2 次
- CodeMMLU: A Multi-Task Benchmark for Assessing Code Understanding & Reasoning Capabilities of CodeLLMsDung Manh Nguyen, Thang Chau Phan, Nam Le Hai, Tien-Thong Doan 等ICLR 2025
它引用的顶会 Paper19
- CoCoNuT: combining context-aware neural translation models using ensemble for program repairThibaud Lutellier, Hung Viet Pham, Lawrence Pang, Yitong Li 等ISSTA 2020 · 被引用 325 次
- Automated Program Repair in the Era of Large Pre-trained Language ModelsChunqiu Steven Xia, Yuxiang Wei, Lingming ZhangICSE 2023 · 被引用 321 次
- CURE: Code-Aware Neural Machine Translation for Automatic Program RepairNan Jiang, Thibaud Lutellier, Lin TanICSE 2021 · 被引用 267 次
- Less training, more repairing please: revisiting automated program repair via zero-shot learningChunqiu Steven Xia, Lingming ZhangFSE 2022 · 被引用 223 次
- A syntax-guided edit decoder for neural program repairQihao Zhu, Zeyu Sun, Yuan-an Xiao, Wenjie Zhang 等FSE 2021 · 被引用 214 次
相关 Paper
- Defects4C: Benchmarking Large Language Model Repair Capability with C/C++ BugsJian Wang, Xiaofei Xie, Qiang Hu, Shangqing Liu 等ASE 2025
- One Size Does Not Fit All: Multi-granularity Patch Generation for Better Automated Program RepairBo Lin, Shangwen Wang, Ming Wen, Liqian Chen 等ISSTA 2024 · 被引用 10 次
- PReMM: LLM-Based Program Repair for Multi-method Bugs via Divide and ConquerLinna Xie, Zhong Li, Yu Pei, Zhongzhen Wen 等OOPSLA 2025 · 被引用 1 次
- How Effective Are Neural Networks for Fixing Security VulnerabilitiesYi Wu, Nan Jiang, Hung Viet Pham, Thibaud Lutellier 等ISSTA 2023 · 被引用 86 次
- A Large-Scale Empirical Review of Patch Correctness Checking ApproachesJun Yang, Yuehan Wang, Yiling Lou, Ming Wen 等FSE 2023 · 被引用 11 次
