Lune

ISSTA2026Top-tier venue

Evaluating the Impact of Explainable AI on Trust in AI-Assisted Code Review

Zhenhan Gao, Marvin Muñoz Barón, Umm-e Habiba, Daniel Graziotin, Stefan Wagner

2026Year

Abstract

Background: The use of large language models (LLMs) for automated code review has brought significant change to a time-consuming part of software engineering. Prior work has shown that LLM-based code tools can improve code quality and enable more robust software development processes. As the tools get more powerful, the explanations behind their decisions remain hard to understand. Developers struggle to assess the validity of LLM-generated code reviews, making it difficult to gauge how much trust they should place in them. While the application of automated code review with LLMs has been extensively investigated, the inclusion of Explainable AI (XAI) for transparency in code reviews and its impact on trust are yet to be explored. Objective: We aim to address this research gap by studying the influence of XAI on the trust of software developers in AI-assisted code reviews. Method: We conducted a within-subjects user study with 34 participants from diverse programming backgrounds, comparing three experimental LLM-based automated code review systems with varying levels of XAI support: Condition A (detailed explanation and review feedback), Condition B (review feedback only), and Condition C (no explanations). Participants were shown a series of real-world code change requests along with the AI-generated code reviews. During the study, we measured trust perceptions for each system using a questionnaire, agreement with the AI recommendation, the reasoning for accepting or rejecting the code change, and the time taken to review the code change. Results: Our quantitative results show that the level of explanation significantly influences both the level of trust of software developers and their agreement with AI recommendations, but in different ways. Full explanations (Condition A) yield the highest perceived trust (M = 3.99/5) but not the highest agreement with AI recommendations, whereas moderate explanations (Condition B) achieve the highest agreement with AI (89.22%). This could suggest that more explanations prompt developers to question AI recommendations more frequently. In contrast, providing no explanations (Condition C) results in the lowest levels of trust and agreement. We also find that the level of explanation did not significantly impact the time taken to accept or reject a code change. Across all conditions, the most commonly cited reasons for code change decisions were changes in code readability and the correctness of the implementation. Conclusion: Overall, these findings indicate that incorporating XAI into the code review process significantly changes the trust perceptions and agreement with AI recommendations for software developers. These results provide insights for the design and evaluation of trustworthy AI-based code review systems, and support researchers in the design of studies on the human factors of AI-assisted software development.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 697d23b6-dda6-4d01-86d3-001834b4d88a

Builds on2

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines