Automated Patch Correctness Assessment: How Far are We?
Shangwen Wang, Ming Wen, Bo Lin, Hongjun Wu, Yihao Qin, Deqing Zou, Xiaoguang Mao, Hai Jin
摘要
Test-based automated program repair (APR) has attracted huge attention from both industry and academia. Despite the significant progress made in recent studies, the overfitting problem (i.e., the generated patch is plausible but overfitting) is still a major and long-standing challenge. Therefore, plenty of techniques have been proposed to assess the correctness of patches either in the patch generation phase or in the evaluation of APR techniques. However, the effectiveness of existing techniques has not been systematically compared and little is known to their advantages and disadvantages. To fill this gap, we performed a large-scale empirical study in this paper. Specifically, we systematically investigated the effectiveness of existing automated patch correctness assessment techniques, including both static and dynamic ones, based on 902 patches automatically generated by 21 APR tools from 4 different categories. Our empirical study revealed the following major findings: (1) static code features with respect to patch syntax and semantics are generally effective in differentiating overfitting patches over correct ones; (2) dynamic techniques can generally achieve high precision while heuristics based on static code features are more effective towards recall; (3) existing techniques are more effective towards certain projects and types of APR techniques while less effective to the others; (4) existing techniques are highly complementary to each other. For instance, a single technique can only detect at most 53.5% of the overfitting patches while 93.3% of them can be detected by at least one technique when the oracle information is available. Based on our findings, we designed an integration strategy to first integrate static code features via learning, and then combine with others by the majority voting strategy. Our experiments show that the strategy can enhance the performance of existing patch correctness assessment techniques significantly.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper26
- Automated Repair of Programs from Large Language ModelsZhiyu Fan, Xiang Gao, Martin Mirchev, Abhik Roychoudhury 等ICSE 2023 · 被引用 213 次
- Neural Program Repair with Execution-based BackpropagationHe Ye, Matias Martinez, Martin MonperrusICSE 2022 · 被引用 146 次
- Baldur: Whole-Proof Generation and Repair with Large Language ModelsEmily First, Markus N. Rabe, Talia Ringer, Yuriy BrunFSE 2023 · 被引用 89 次
- Evaluating Representation Learning of Code Changes for Predicting Patch Correctness in Program RepairHaoye Tian, Kui Liu, Abdoul Kader Kaboré, Anil Koyuncu 等ASE 2020 · 被引用 81 次
- CCT5: A Code-Change-Oriented Pre-trained ModelBo Lin, Shangwen Wang, Zhongxin Liu, Yepang Liu 等FSE 2023 · 被引用 69 次
它引用的顶会 Paper4
- On the efficiency of test suite based program repair: A Systematic Assessment of 16 Automated Repair Systems for Java ProgramsKui Liu, Shangwen Wang, Anil Koyuncu, Kisub Kim 等ICSE 2020 · 被引用 116 次
- Can automated program repair refine fault localization? a unified debugging approachYiling Lou, Ali Ghanbari, Xia Li, Lingming Zhang 等ISSTA 2020 · 被引用 99 次
- Evaluating Representation Learning of Code Changes for Predicting Patch Correctness in Program RepairHaoye Tian, Kui Liu, Abdoul Kader Kaboré, Anil Koyuncu 等ASE 2020 · 被引用 81 次
- On the Effectiveness of Unified Debugging: An Extensive Study on 16 Program Repair SystemsSamuel Benton, Xia Li, Yiling Lou, Lingming ZhangASE 2020 · 被引用 35 次
相关 Paper
- A Large-Scale Empirical Review of Patch Correctness Checking ApproachesJun Yang, Yuehan Wang, Yiling Lou, Ming Wen 等FSE 2023 · 被引用 11 次
- Enhancing APR with PRISM: A Semantic-Based Approach to Overfitting Patch DetectionDowon Song, Hakjoo OhOOPSLA 2025 · 被引用 2 次
- Patch correctness assessment in automated program repair based on the impact of patches on production and test codeAli Ghanbari, Andrian MarcusISSTA 2022 · 被引用 26 次
- Practical Program Repair via Preference-based Ensemble StrategyWenkang Zhong, Chuanyi Li, Kui Liu, Tongtong Xu 等ICSE 2024 · 被引用 8 次
- Towards Boosting Patch Execution On-the-FlySamuel Benton, Yuntong Xie, Lan Lu, Mengshi Zhang 等ICSE 2022 · 被引用 10 次
