Lune

ISSTA2026顶会

E4R-Reviewer: Effective and Explainable Automated Code Review via End-to-End Reasoning-Guided Alignment

Yifei Liu, Xizhi Hou, Li Yang, Huan Liu, Chen Zhu, Fengjun Zhang, Chun Zuo

2026年份

摘要

Code review is a key practice for ensuring software quality and maintainability. Despite progress in Automated Code Review (ACR), existing methods face two core challenges: (1) Isolated Task Modeling. Current approaches often model and optimize subtasks in ACR independently, ignoring the inherent logical order and internal dependencies among them which affects the effectiveness of ACR. (2) Lack of Explainability. At the task level, the absence of explanatory information in review comments increases developers’ cognitive load; at the model level, the black-box nature fundamentally undermines developer trust. To address these challenges, we propose E4R-Reviewer, which improves the E ffec­tiveness and E xplain­ability of ACR through E nd-to- E nd R easoning-guided alignment. For effectiveness , E4R-Reviewer unifies multiple fine-grained ACR subtasks into a single end-to-end reasoning process, enabling cross-task knowledge sharing and allowing the model to explicitly complete a reasoning chain that covers quality estimation, issue localization, issue classification, issue description, fix suggestion, and code refinement in one generation. Meanwhile, we adopt a Group Relative Policy Optimization (GRPO)-based reinforcement-learning alignment, treating the reasoning steps as optimizable intermediate objectives. We design subtask-specific rewards and integrate them via curriculum-inspired, multi-stage reward fusion that follows the real-world review workflow. For explainability , E4R-Reviewer produces reasoning process and structured review results covering all fine-grained ACR subtasks, improving the transparency and explainability of the review results. Extensive evaluations on public, real-world datasets demonstrate that E4R-Reviewer significantly outperforms existing methods and achieves state-of-the-art performance: a 74.61% F1-score in quality estimation and +22.96% CodeBLEU in code refinement. Furthermore, Large Language Model (LLM) and human evaluation further confirm the superiority of E4R-Reviewer in terms of effectiveness and explainability.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖