CCT5: A Code-Change-Oriented Pre-trained Model
Bo Lin, Shangwen Wang, Zhongxin Liu, Yepang Liu, Xin Xia, Xiaoguang Mao
Abstract
Software is constantly changing, requiring developers to perform several derived tasks in a timely manner, such as writing a description for the intention of the code change, or identifying the defect-prone code changes. Considering that the cost of dealing with these tasks can account for a large proportion (typically around 70 percent) of the total development expenditure, automating such processes will significantly lighten the burdens of developers. To achieve such a target, existing approaches mainly rely on training deep learning models from scratch or fine-tuning existing pre-trained models on such tasks, both of which have weaknesses. Specifically, the former uses comparatively small-scale labelled data for training, making it difficult to learn and exploit the domain knowledge of programming language hidden in the large-amount unlabelled code in the wild; the latter is hard to fully leverage the learned knowledge of the pre-trained model, as existing pre-trained models are designed to encode a single code snippet rather than a code change (the difference between two code snippets). We propose to pre-train a model specially designed for code changes to better support developers in software maintenance. To this end, we first collect a large-scale dataset containing 1.5M+ pairwise data of code changes and commit messages. Based on these data, we curate five different tasks for pre-training, which equip the model with diverse domain knowledge about code changes. We fine-tune the pre-trained model, CCT5, on three widely-studied tasks incurred by code changes and two tasks specific to the code review process. Results show that CCT5 outperforms both conventional deep learning approaches and existing pre-trained models on these tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 843939c7-52d5-46d2-9330-75bbe414c2b9Cited by top-tier papers17
- CoEdPilot: Recommending Code Edits with Learned Prior Edit Relevance, Project-wise Awareness, and Interactive NatureChenyan Liu, Yufan Cai, Yun Lin, Yuhuan Huang et al.ISSTA 2024 · 7 citations
- Understanding Code Changes Practically with Small-Scale Language ModelsCong Li, Zhaogui Xu, Peng Di, Dongxia Wang et al.ASE 2024 · 4 citations
- EfficientEdit: Accelerating Code Editing via Edit-Oriented Speculative DecodingPeiding Wang, Li Zhang, Fang Liu, Yinghao Zhu et al.ASE 2025 · 3 citations
- Instruct or Interact? Exploring and Eliciting LLMs' Capability in Code Snippet Adaptation Through Prompt EngineeringTanghaoran Zhang, Yue Yu, Xinjun Mao, Shangwen Wang et al.ICSE 2025 · 3 citations
- Spotting Code Mutation for Predictive Mutation TestingYifan Zhao, Yizhou Chen, Zeyu Sun, Qingyuan Liang et al.ASE 2024 · 2 citations
Builds on17
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng et al.ICLR 2021 · 1,644 citations
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 1,224 citations
- VulRepair: a T5-based automated software vulnerability repairMichael Fu, Chakkrit Tantithamthavorn, Trung Le, Van Nguyen et al.FSE 2022 · 206 citations
- Automating code review activities by large-scale pre-trainingZhiyu Li, Shuai Lu, Daya Guo, Nan Duan et al.FSE 2022 · 195 citations
- CC2Vec: distributed representations of code changesThong Hoang, Hong Jin Kang, David Lo, Julia LawallICSE 2020 · 169 citations
Related papers
- Using Pre-Trained Models to Boost Code Review AutomationRosalia Tufano, Simone Masiero, Antonio Mastropaolo, Luca Pascarella et al.ICSE 2022 · 149 citations
- Towards Automating Code Review ActivitiesRosalia Tufano, Luca Pascarella, Michele Tufano, Denys Poshyvanyk et al.ICSE 2021 · 4 citations
- CoditT5: Pretraining for Source Code and Natural Language EditingJiyang Zhang, Sheena Panthaplackel, Pengyu Nie, Junyi Jessy Li et al.ASE 2022 · 81 citations
- Deep Just-In-Time Inconsistency Detection Between Comments and Source CodeSheena Panthaplackel, Junyi Jessy Li, Milos Gligoric, Raymond J. MooneyAAAI 2021 · 62 citations
- Studying the Usage of Text-To-Text Transfer Transformer to Support Code-Related TasksAntonio Mastropaolo, Simone Scalabrino, Nathan Cooper, David Nader-Palacio et al.ICSE 2021 · 9 citations
