Evaluating Representation Learning of Code Changes for Predicting Patch Correctness in Program Repair
Haoye Tian, Kui Liu, Abdoul Kader Kaboré, Anil Koyuncu, Li Li, Jacques Klein, Tegawendé F. Bissyandé
Abstract
A large body of the literature of automated program repair develops approaches where patches are generated to be validated against an oracle (e.g., a test suite). Because such an oracle can be imperfect, the generated patches, although validated by the oracle, may actually be incorrect. While the state of the art explore research directions that require dynamic information or that rely on manually-crafted heuristics, we study the benefit of learning code representations in order to learn deep features that may encode the properties of patch correctness. Our empirical work mainly investigates different representation learning approaches for code changes to derive embeddings that are amenable to similarity computations. We report on findings based on embeddings produced by pre-trained and re-trained neural networks. Experimental results demonstrate the potential of embeddings to empower learning algorithms in reasoning about patch correctness: a machine learning predictor with BERT transformer-based embeddings associated with logistic regression yielded an AUC value of about 0.8 in the prediction of patch correctness on a deduplicated dataset of 1000 labeled patches. Our investigations show that learned representations can lead to reasonable performance when comparing against the state-of-the-art, PATCH-SIM, which relies on dynamic information. These representations may further be complementary to features that were carefully (manually) engineered in the literature. CCS CONCEPTS • Software and its engineering → Software testing and debugging.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bdccfe4f-5ae9-4396-a341-061b431ae051Cited by top-tier papers25
- CURE: Code-Aware Neural Machine Translation for Automatic Program RepairNan Jiang, Thibaud Lutellier, Lin TanICSE 2021 · 267 citations
- Neural Program Repair with Execution-based BackpropagationHe Ye, Matias Martinez, Martin MonperrusICSE 2022 · 146 citations
- Baldur: Whole-Proof Generation and Repair with Large Language ModelsEmily First, Markus N. Rabe, Talia Ringer, Yuriy BrunFSE 2023 · 89 citations
- Automated Patch Correctness Assessment: How Far are We?Shangwen Wang, Ming Wen, Bo Lin, Hongjun Wu et al.ASE 2020 · 77 citations
- DeepCVA: Automated Commit-level Vulnerability Assessment with Deep Multi-task LearningTriet Huynh Minh Le, David Hin, Roland Croft, Muhammad Ali BabarASE 2021 · 62 citations
Builds on4
- Order Matters: Semantic-Aware Neural Networks for Binary Code Similarity DetectionZeping Yu, Rui Cao, Qiyi Tang, Sen Nie et al.AAAI 2020 · 265 citations
- CC2Vec: distributed representations of code changesThong Hoang, Hong Jin Kang, David Lo, Julia LawallICSE 2020 · 169 citations
- On the efficiency of test suite based program repair: A Systematic Assessment of 16 Automated Repair Systems for Java ProgramsKui Liu, Shangwen Wang, Anil Koyuncu, Kisub Kim et al.ICSE 2020 · 116 citations
- Automated Patch Correctness Assessment: How Far are We?Shangwen Wang, Ming Wen, Bo Lin, Hongjun Wu et al.ASE 2020 · 77 citations
Related papers
- Is this Change the Answer to that Problem?: Correlating Descriptions of Bug and Code Changes for Evaluating Patch CorrectnessHaoye Tian, Xunzhu Tang, Andrew Habib, Shangwen Wang et al.ASE 2022 · 11 citations
- Less training, more repairing please: revisiting automated program repair via zero-shot learningChunqiu Steven Xia, Lingming ZhangFSE 2022 · 223 citations
- Learning and Evaluating Contextual Embedding of Source CodeAditya Kanade, Petros Maniatis, Gogul Balakrishnan, Kensen ShiICML 2020 · 438 citations
- CCRep: Learning Code Change Representations via Pre-Trained Code Model and Query BackZhongxin Liu, Zhijie Tang, Xin Xia, Xiaohu YangICSE 2023 · 25 citations
- RAP-Gen: Retrieval-Augmented Patch Generation with CodeT5 for Automatic Program RepairWeishi Wang, Yue Wang, Shafiq Joty, Steven C. H. HoiFSE 2023 · 84 citations
