Deep Learning-based Code Reviews: A Paradigm Shift or a Double-Edged Sword?
Rosalia Tufano, Alberto Martin-Lopez, Ahmad Tayeb, Ozren Dabic, Sonia Haiduc, Gabriele Bavota
Abstract
Several techniques have been proposed to (partially) automate code review. Early support consisted in recommending the most suited reviewer for a given change or in prioritizing the review tasks. With the advent of deep learning in software engineering, the level of automation has been pushed to new heights, with approaches able to provide feedback on source code in natural language as a human reviewer would do. Also, recent work documented open source projects adopting Large Language Models (LLMs) as co-reviewers. Although the research in this field is very active, little is known about the actual impact of including automatically generated code reviews in the code review process. While there are many aspects worth investigating (e.g., is knowledge transfer between developers affected?), in this work we focus on three of them: (i) review quality, i.e., the reviewer's ability to identify issues in the code; (ii) review cost, i.e., the time spent reviewing the code; and (iii) reviewer's confidence, i.e., how confident is the reviewer about the provided feedback. We run a controlled experiment with 29 professional developers who reviewed different programs with/without the support of an automatically generated code review. During the experiment we monitored the reviewers' activities, for over 50 hours of recorded code reviews. We show that reviewers consider valid most of the issues automatically identified by the LLM and that the availability of an automated review as a starting point strongly influences their behavior: Reviewers tend to focus on the code locations indicated by the LLM rather than searching for additional issues in other parts of the code. The reviewers who started from an automated review identified a higher number of low-severity issues while, however, not identifying more high-severity issues as compared to a completely manual process. Finally, the automated support did not result in saved time and did not increase the reviewers' confidence.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a25e1b85-a565-47b0-8e73-078c5e690db3Builds on14
- Using an LLM to Help With Code UnderstandingDaye Nam, Andrew Macvean, Vincent J. Hellendoorn, Bogdan Vasilescu et al.ICSE 2024 · 264 citations
- Automating code review activities by large-scale pre-trainingZhiyu Li, Shuai Lu, Daya Guo, Nan Duan et al.FSE 2022 · 195 citations
- Using Pre-Trained Models to Boost Code Review AutomationRosalia Tufano, Simone Masiero, Antonio Mastropaolo, Luca Pascarella et al.ICSE 2022 · 149 citations
- Make LLM a Testing Expert: Bringing Human-like Interaction to Mobile GUI Testing via Functionality-aware DecisionsZhe Liu, Chunyang Chen, Junjie Wang, Mengzhuo Chen et al.ICSE 2024 · 81 citations
- LLMParser: An Exploratory Study on Using Large Language Models for Log ParsingZeyang Ma, An Ran Chen, Dong Jae Kim, Tse-Hsun Chen et al.ICSE 2024 · 72 citations
Related papers
- Towards Automating Code Review ActivitiesRosalia Tufano, Luca Pascarella, Michele Tufano, Denys Poshyvanyk et al.ICSE 2021 · 4 citations
- Evaluating the Impact of Explainable AI on Trust in AI-Assisted Code ReviewZhenhan Gao, Marvin Muñoz Barón, Umm-e Habiba, Daniel Graziotin et al.ISSTA 2026
- LAURA: Enhancing Code Review Generation with Context-Enriched Retrieval-Augmented LLMYuxin Zhang, Yuxia Zhang, Zeyu Sun, Yanjie Jiang et al.ASE 2025 · 8 citations
- Intention is All you Need: Refining your Code from your IntentionQi Guo, Xiaofei Xie, Shangqing Liu, Ming Hu et al.ICSE 2025 · 7 citations
- CORE: Resolving Code Quality Issues using LLMsNalin Wadhwa, Jui Pradhan, Atharv Sonwane, Surya Prakash Sahu et al.FSE 2024 · 32 citations
