E4R-Reviewer: Effective and Explainable Automated Code Review via End-to-End Reasoning-Guided Alignment
Yifei Liu, Xizhi Hou, Li Yang, Huan Liu, Chen Zhu, Fengjun Zhang, Chun Zuo
Abstract
Code review is a key practice for ensuring software quality and maintainability. Despite progress in Automated Code Review (ACR), existing methods face two core challenges: (1) Isolated Task Modeling. Current approaches often model and optimize subtasks in ACR independently, ignoring the inherent logical order and internal dependencies among them which affects the effectiveness of ACR. (2) Lack of Explainability. At the task level, the absence of explanatory information in review comments increases developers’ cognitive load; at the model level, the black-box nature fundamentally undermines developer trust. To address these challenges, we propose E4R-Reviewer, which improves the E ffectiveness and E xplainability of ACR through E nd-to- E nd R easoning-guided alignment. For effectiveness , E4R-Reviewer unifies multiple fine-grained ACR subtasks into a single end-to-end reasoning process, enabling cross-task knowledge sharing and allowing the model to explicitly complete a reasoning chain that covers quality estimation, issue localization, issue classification, issue description, fix suggestion, and code refinement in one generation. Meanwhile, we adopt a Group Relative Policy Optimization (GRPO)-based reinforcement-learning alignment, treating the reasoning steps as optimizable intermediate objectives. We design subtask-specific rewards and integrate them via curriculum-inspired, multi-stage reward fusion that follows the real-world review workflow. For explainability , E4R-Reviewer produces reasoning process and structured review results covering all fine-grained ACR subtasks, improving the transparency and explainability of the review results. Extensive evaluations on public, real-world datasets demonstrate that E4R-Reviewer significantly outperforms existing methods and achieves state-of-the-art performance: a 74.61% F1-score in quality estimation and +22.96% CodeBLEU in code refinement. Furthermore, Large Language Model (LLM) and human evaluation further confirm the superiority of E4R-Reviewer in terms of effectiveness and explainability.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 25ed4475-98a0-467f-9f6b-5fdb3f192086Related papers
- Exploring the Potential of ChatGPT in Automated Code Refinement: An Empirical StudyQi Guo, Junming Cao, Xiaofei Xie, Shangqing Liu et al.ICSE 2024 · 107 citations
- Improving the Learning of Code Review Successive Tasks with Cross-Task Knowledge DistillationOussama Ben Sghaier, Houari A. SahraouiFSE 2024 · 9 citations
- Automating code review activities by large-scale pre-trainingZhiyu Li, Shuai Lu, Daya Guo, Nan Duan et al.FSE 2022 · 195 citations
- CodeAgent: Autonomous Communicative Agents for Code ReviewXunzhu Tang, Kisub Kim, Yewei Song, Cedric Lothritz et al.EMNLP 2024 · 8 citations
- VisualScore: Learning Holistic Visual Quality Scores via Multi-Task ReasoningYiting Lu, Fengbin Guan, Yixin Gao, Yan Zhong et al.ICML 2026
