Lune

ISSTA2026顶会

Evaluating the Impact of Explainable AI on Trust in AI-Assisted Code Review

Zhenhan Gao, Marvin Muñoz Barón, Umm-e Habiba, Daniel Graziotin, Stefan Wagner

2026年份

摘要

Background: The use of large language models (LLMs) for automated code review has brought significant change to a time-consuming part of software engineering. Prior work has shown that LLM-based code tools can improve code quality and enable more robust software development processes. As the tools get more powerful, the explanations behind their decisions remain hard to understand. Developers struggle to assess the validity of LLM-generated code reviews, making it difficult to gauge how much trust they should place in them. While the application of automated code review with LLMs has been extensively investigated, the inclusion of Explainable AI (XAI) for transparency in code reviews and its impact on trust are yet to be explored. Objective: We aim to address this research gap by studying the influence of XAI on the trust of software developers in AI-assisted code reviews. Method: We conducted a within-subjects user study with 34 participants from diverse programming backgrounds, comparing three experimental LLM-based automated code review systems with varying levels of XAI support: Condition A (detailed explanation and review feedback), Condition B (review feedback only), and Condition C (no explanations). Participants were shown a series of real-world code change requests along with the AI-generated code reviews. During the study, we measured trust perceptions for each system using a questionnaire, agreement with the AI recommendation, the reasoning for accepting or rejecting the code change, and the time taken to review the code change. Results: Our quantitative results show that the level of explanation significantly influences both the level of trust of software developers and their agreement with AI recommendations, but in different ways. Full explanations (Condition A) yield the highest perceived trust (M = 3.99/5) but not the highest agreement with AI recommendations, whereas moderate explanations (Condition B) achieve the highest agreement with AI (89.22%). This could suggest that more explanations prompt developers to question AI recommendations more frequently. In contrast, providing no explanations (Condition C) results in the lowest levels of trust and agreement. We also find that the level of explanation did not significantly impact the time taken to accept or reject a code change. Across all conditions, the most commonly cited reasons for code change decisions were changes in code readability and the correctness of the implementation. Conclusion: Overall, these findings indicate that incorporating XAI into the code review process significantly changes the trust perceptions and agreement with AI recommendations for software developers. These results provide insights for the design and evaluation of trustworthy AI-based code review systems, and support researchers in the design of studies on the human factors of AI-assisted software development.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper2

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖