Automating code review activities by large-scale pre-training
Zhiyu Li, Shuai Lu, Daya Guo, Nan Duan, Shailesh Jannu, Grant Jenks, Deep Majumder, Jared Green, Alexey Svyatkovskiy, Shengyu Fu, Neel Sundaresan
摘要
Code review is an essential part to software development lifecycle since it aims at guaranteeing the quality of codes. Modern code review activities necessitate developers viewing, understanding and even running the programs to assess logic, functionality, latency, style and other factors. It turns out that developers have to spend far too much time reviewing the code of their peers. Accordingly, it is in significant demand to automate the code review process. In this research, we focus on utilizing pre-training techniques for the tasks in the code review scenario. We collect a large-scale dataset of real-world code changes and code reviews from open-source projects in nine of the most popular programming languages. To better understand code diffs and reviews, we propose CodeReviewer, a pre-trained model that utilizes four pre-training tasks tailored specifically for the code review scenario. To evaluate our model, we focus on three key tasks related to code review activities, including code change quality estimation, review comment generation and code refinement. Furthermore, we establish a high-quality benchmark dataset based on our collected data for these three tasks and conduct comprehensive experiments on it. The experimental results demonstrate that our model outperforms the previous state-of-the-art pre-training approaches in all tasks. Further analysis show that our proposed pre-training tasks and the multilingual pre-training dataset benefit the model on the understanding of code changes and reviews.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper40
- Exploring the Potential of ChatGPT in Automated Code Refinement: An Empirical StudyQi Guo, Junming Cao, Xiaofei Xie, Shangqing Liu 等ICSE 2024 · 被引用 107 次
- NExT: Teaching Large Language Models to Reason about Code ExecutionAnsong Ni, Miltiadis Allamanis, Arman Cohan, Yinlin Deng 等ICML 2024 · 被引用 73 次
- CCT5: A Code-Change-Oriented Pre-trained ModelBo Lin, Shangwen Wang, Zhongxin Liu, Yepang Liu 等FSE 2023 · 被引用 69 次
- Gamma: Revisiting Template-Based Automated Program Repair Via Mask PredictionQuanjun Zhang, Chunrong Fang, Tongke Zhang, Bowen Yu 等ASE 2023 · 被引用 44 次
- Multilingual Code Co-evolution using Large Language ModelsJiyang Zhang, Pengyu Nie, Junyi Jessy Li, Milos GligoricFSE 2023 · 被引用 34 次
它引用的顶会 Paper7
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng 等ICLR 2021 · 被引用 1,644 次
- Learning and Evaluating Contextual Embedding of Source CodeAditya Kanade, Petros Maniatis, Gogul Balakrishnan, Kensen ShiICML 2020 · 被引用 438 次
- Using Pre-Trained Models to Boost Code Review AutomationRosalia Tufano, Simone Masiero, Antonio Mastropaolo, Luca Pascarella 等ICSE 2022 · 被引用 149 次
- Multilingual Code Snippets Training for Program TranslationMing Zhu, Karthik Suresh, Chandan K. ReddyAAAI 2022 · 被引用 72 次
相关 Paper
- Towards Automating Code Review ActivitiesRosalia Tufano, Luca Pascarella, Michele Tufano, Denys Poshyvanyk 等ICSE 2021 · 被引用 4 次
- CoditT5: Pretraining for Source Code and Natural Language EditingJiyang Zhang, Sheena Panthaplackel, Pengyu Nie, Junyi Jessy Li 等ASE 2022 · 被引用 81 次
- Improving the Learning of Code Review Successive Tasks with Cross-Task Knowledge DistillationOussama Ben Sghaier, Houari A. SahraouiFSE 2024 · 被引用 9 次
- AUGER: automatically generating review comments with pre-training modelsLingwei Li, Li Yang, Huaxi Jiang, Jun Yan 等FSE 2022 · 被引用 56 次
- CommentFinder: a simpler, faster, more accurate code review comments recommendationYang Hong, Chakkrit Tantithamthavorn, Patanamon Thongtanunam, Aldeida AletiFSE 2022 · 被引用 57 次
