Improving the Learning of Code Review Successive Tasks with Cross-Task Knowledge Distillation
Oussama Ben Sghaier, Houari A. Sahraoui
摘要
Code review is a fundamental process in software development that plays a pivotal role in ensuring code quality and reducing the likelihood of errors and bugs. However, code review can be complex, subjective, and time-consuming. Quality estimation , comment generation , and code refinement constitute the three key tasks of this process, and their automation has traditionally been addressed separately in the literature using different approaches. In particular, recent efforts have focused on fine-tuning pre-trained language models to aid in code review tasks, with each task being considered in isolation. We believe that these tasks are interconnected, and their fine-tuning should consider this interconnection. In this paper, we introduce a novel deep-learning architecture, named DISCOREV, which employs cross-task knowledge distillation to address these tasks simultaneously. In our approach, we utilize a cascade of models to enhance both comment generation and code refinement models. The fine-tuning of the comment generation model is guided by the code refinement model, while the fine-tuning of the code refinement model is guided by the quality estimation model. We implement this guidance using two strategies: a feedback-based learning objective and an embedding alignment objective. We evaluate DISCOREV by comparing it to state-of-the-art methods based on independent training and fine-tuning. Our results show that our approach generates better review comments, as measured by the BLEU score, as well as more accurate code refinement according to the CodeBLEU score.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- SWR-Bench: Assessing LLM Performance in Real-World Code Review Comment GenerationZhengran Zeng, Ruikai Shi, Keke Han, Yixin Li 等FSE 2026
- Towards Practical Defect-Focused Automated Code ReviewJunyi Lu, Lili Jiang, Xiaojia Li, Jianbing Fang 等ICML 2025
- SeRe: A Security-Related Code Review Dataset Aligned with Real-World Review ActivitiesZixiao Zhao, Yanjie Jiang, Hui Liu, Kui Liu 等ICSE 2026
它引用的顶会 Paper8
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 被引用 1,224 次
- Automating code review activities by large-scale pre-trainingZhiyu Li, Shuai Lu, Daya Guo, Nan Duan 等FSE 2022 · 被引用 195 次
- Using Pre-Trained Models to Boost Code Review AutomationRosalia Tufano, Simone Masiero, Antonio Mastropaolo, Luca Pascarella 等ICSE 2022 · 被引用 149 次
- Multilingual Code Snippets Training for Program TranslationMing Zhu, Karthik Suresh, Chandan K. ReddyAAAI 2022 · 被引用 72 次
- Cross-Task Knowledge Distillation in Multi-Task RecommendationChenxiao Yang, Junwei Pan, Xiaofeng Gao, Tingyu Jiang 等AAAI 2022 · 被引用 58 次
相关 Paper
- E4R-Reviewer: Effective and Explainable Automated Code Review via End-to-End Reasoning-Guided AlignmentYifei Liu, Xizhi Hou, Li Yang, Huan Liu 等ISSTA 2026
- CoditT5: Pretraining for Source Code and Natural Language EditingJiyang Zhang, Sheena Panthaplackel, Pengyu Nie, Junyi Jessy Li 等ASE 2022 · 被引用 81 次
- Intention is All you Need: Refining your Code from your IntentionQi Guo, Xiaofei Xie, Shangqing Liu, Ming Hu 等ICSE 2025 · 被引用 7 次
- Studying the Usage of Text-To-Text Transfer Transformer to Support Code-Related TasksAntonio Mastropaolo, Simone Scalabrino, Nathan Cooper, David Nader-Palacio 等ICSE 2021 · 被引用 9 次
- Towards Automating Code Review ActivitiesRosalia Tufano, Luca Pascarella, Michele Tufano, Denys Poshyvanyk 等ICSE 2021 · 被引用 4 次
