CCRep: Learning Code Change Representations via Pre-Trained Code Model and Query Back
Zhongxin Liu, Zhijie Tang, Xin Xia, Xiaohu Yang
Abstract
Representing code changes as numeric feature vectors, i.e., code change representations, is usually an essential step to automate many software engineering tasks related to code changes, e.g., commit message generation and just-in-time defect prediction. Intuitively, the quality of code change representations is crucial for the effectiveness of automated approaches. Prior work on code changes usually designs and evaluates code change representation approaches for a specific task, and little work has investigated code change encoders that can be used and jointly trained on various tasks. To fill this gap, this work proposes a novel Code Change Representation learning approach named CCRep, which can learn to encode code changes as feature vectors for diverse downstream tasks. Specifically, CCRep regards a code change as the combination of its before-change and after-change code, leverages a pre-trained code model to obtain high-quality contextual embeddings of code, and uses a novel mechanism named query back to extract and encode the changed code fragments and make them explicitly interact with the whole code change. To evaluate CCRep and demonstrate its applicability to diverse code-change-related tasks, we apply it to three tasks: commit message generation, patch correctness assessment, and just-in-time defect prediction. Experimental results show that CCRep outperforms the state-of-the-art techniques on each task.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8d31f7b0-1e58-4441-8a79-708b29e8cff9Cited by top-tier papers1
Ask how each one uses itBuilds on16
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 1,224 citations
- CC2Vec: distributed representations of code changesThong Hoang, Hong Jin Kang, David Lo, Julia LawallICSE 2020 · 169 citations
- Multi-task Learning based Pre-trained Language Model for Code CompletionFang Liu, Ge Li, Yunfei Zhao, Zhi JinASE 2020 · 162 citations
- Big code != big vocabulary: open-vocabulary models for source codeRafael-Michael Karampatsis, Hlib Babii, Romain Robbes, Charles Sutton et al.ICSE 2020 · 140 citations
- Traceability Transformed: Generating more Accurate Links with Pre-Trained BERT ModelsJinfeng Lin, Yalin Liu, Qingkai Zeng, Meng Jiang et al.ICSE 2021 · 124 citations
Related papers
- COME: Commit Message Generation with Modification EmbeddingYichen He, Liran Wang, Kaiyi Wang, Yupeng Zhang et al.ISSTA 2023 · 14 citations
- CCT5: A Code-Change-Oriented Pre-trained ModelBo Lin, Shangwen Wang, Zhongxin Liu, Yepang Liu et al.FSE 2023 · 69 citations
- Learning to Update Natural Language Comments Based on Code ChangesSheena Panthaplackel, Pengyu Nie, Milos Gligoric, Junyi Jessy Li et al.ACL 2020 · 1 citation
- Automating Just-In-Time Comment UpdatingZhongxin Liu, Xin Xia, Meng Yan, Shanping LiASE 2020 · 46 citations
- Jointly Learning to Repair Code and Generate Commit MessageJiaqi Bai, Long Zhou, Ambrosio Blanco, Shujie Liu et al.EMNLP 2021 · 5 citations
