E4R-Reviewer: Effective and Explainable Automated Code Review via End-to-End Reasoning-Guided Alignment
Yifei Liu, Xizhi Hou, Li Yang, Huan Liu, Chen Zhu, Fengjun Zhang, Chun Zuo
摘要
Code review is a key practice for ensuring software quality and maintainability. Despite progress in Automated Code Review (ACR), existing methods face two core challenges: (1) Isolated Task Modeling. Current approaches often model and optimize subtasks in ACR independently, ignoring the inherent logical order and internal dependencies among them which affects the effectiveness of ACR. (2) Lack of Explainability. At the task level, the absence of explanatory information in review comments increases developers’ cognitive load; at the model level, the black-box nature fundamentally undermines developer trust. To address these challenges, we propose E4R-Reviewer, which improves the E ffectiveness and E xplainability of ACR through E nd-to- E nd R easoning-guided alignment. For effectiveness , E4R-Reviewer unifies multiple fine-grained ACR subtasks into a single end-to-end reasoning process, enabling cross-task knowledge sharing and allowing the model to explicitly complete a reasoning chain that covers quality estimation, issue localization, issue classification, issue description, fix suggestion, and code refinement in one generation. Meanwhile, we adopt a Group Relative Policy Optimization (GRPO)-based reinforcement-learning alignment, treating the reasoning steps as optimizable intermediate objectives. We design subtask-specific rewards and integrate them via curriculum-inspired, multi-stage reward fusion that follows the real-world review workflow. For explainability , E4R-Reviewer produces reasoning process and structured review results covering all fine-grained ACR subtasks, improving the transparency and explainability of the review results. Extensive evaluations on public, real-world datasets demonstrate that E4R-Reviewer significantly outperforms existing methods and achieves state-of-the-art performance: a 74.61% F1-score in quality estimation and +22.96% CodeBLEU in code refinement. Furthermore, Large Language Model (LLM) and human evaluation further confirm the superiority of E4R-Reviewer in terms of effectiveness and explainability.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Exploring the Potential of ChatGPT in Automated Code Refinement: An Empirical StudyQi Guo, Junming Cao, Xiaofei Xie, Shangqing Liu 等ICSE 2024 · 被引用 107 次
- Improving the Learning of Code Review Successive Tasks with Cross-Task Knowledge DistillationOussama Ben Sghaier, Houari A. SahraouiFSE 2024 · 被引用 9 次
- Automating code review activities by large-scale pre-trainingZhiyu Li, Shuai Lu, Daya Guo, Nan Duan 等FSE 2022 · 被引用 195 次
- CodeAgent: Autonomous Communicative Agents for Code ReviewXunzhu Tang, Kisub Kim, Yewei Song, Cedric Lothritz 等EMNLP 2024 · 被引用 8 次
- VisualScore: Learning Holistic Visual Quality Scores via Multi-Task ReasoningYiting Lu, Fengbin Guan, Yixin Gao, Yan Zhong 等ICML 2026
