An Empirical Study on Code Review Activity Prediction and Its Impact in Practice
Doriane Olewicki, Sarra Habchi, Bram Adams
摘要
During code reviews, an essential step in software quality assurance, reviewers have the difficult task of understanding and evaluating code changes to validate their quality and prevent introducing faults to the codebase. This is a tedious process where the effort needed is highly dependent on the code submitted, as well as the author’s and the reviewer’s experience, leading to median wait times for review feedback of 15-64 hours. Through an initial user study carried with 29 experts, we found that re-ordering the files changed by a patch within the review environment has potential to improve review quality, as more comments are written (+23%), and participants’ file-level hot-spot precision and recall increases to 53% (+13%) and 28% (+8%), respectively, compared to the alphanumeric ordering. Hence, this paper aims to help code reviewers by predicting which files in a submitted patch need to be (1) commented, (2) revised, or (3) are hot-spots (commented or revised). To predict these tasks, we evaluate two different types of text embeddings (i.e., Bag-of-Words and Large Language Models encoding) and review process features (i.e., code size-based and history-based features). Our empirical study on three open-source and two industrial datasets shows that combining the code embedding and review process features leads to better results than the state-of-the-art approach. For all tasks, F1-scores (median of 40-62%) are significantly better than the state-of-the-art (from +1 to +9%).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper8
- Transformer in TransformerKai Han, An Xiao, Enhua Wu, Jianyuan Guo 等NeurIPS 2021 · 被引用 2,148 次
- Automating code review activities by large-scale pre-trainingZhiyu Li, Shuai Lu, Daya Guo, Nan Duan 等FSE 2022 · 被引用 195 次
- Using Pre-Trained Models to Boost Code Review AutomationRosalia Tufano, Simone Masiero, Antonio Mastropaolo, Luca Pascarella 等ICSE 2022 · 被引用 149 次
- Large Language Models Meet NL2Code: A SurveyDaoguang Zan, Bei Chen, Fengji Zhang, Dianjie Lu 等ACL 2023 · 被引用 104 次
- CommentFinder: a simpler, faster, more accurate code review comments recommendationYang Hong, Chakkrit Tantithamthavorn, Patanamon Thongtanunam, Aldeida AletiFSE 2022 · 被引用 57 次
相关 Paper
- Deep Learning-based Code Reviews: A Paradigm Shift or a Double-Edged Sword?Rosalia Tufano, Alberto Martin-Lopez, Ahmad Tayeb, Ozren Dabic 等ICSE 2025 · 被引用 1 次
- Intention is All you Need: Refining your Code from your IntentionQi Guo, Xiaofei Xie, Shangqing Liu, Ming Hu 等ICSE 2025 · 被引用 7 次
- LAURA: Enhancing Code Review Generation with Context-Enriched Retrieval-Augmented LLMYuxin Zhang, Yuxia Zhang, Zeyu Sun, Yanjie Jiang 等ASE 2025 · 被引用 8 次
- Using nudges to accelerate code reviews at scaleQianhua Shan, David Sukhdeo, Qianying Huang, Seth Rogers 等FSE 2022 · 被引用 17 次
- CORE: Resolving Code Quality Issues using LLMsNalin Wadhwa, Jui Pradhan, Atharv Sonwane, Surya Prakash Sahu 等FSE 2024 · 被引用 32 次
