Evaluating the Impact of Explainable AI on Trust in AI-Assisted Code Review
Zhenhan Gao, Marvin Muñoz Barón, Umm-e Habiba, Daniel Graziotin, Stefan Wagner
Abstract
Background: The use of large language models (LLMs) for automated code review has brought significant change to a time-consuming part of software engineering. Prior work has shown that LLM-based code tools can improve code quality and enable more robust software development processes. As the tools get more powerful, the explanations behind their decisions remain hard to understand. Developers struggle to assess the validity of LLM-generated code reviews, making it difficult to gauge how much trust they should place in them. While the application of automated code review with LLMs has been extensively investigated, the inclusion of Explainable AI (XAI) for transparency in code reviews and its impact on trust are yet to be explored. Objective: We aim to address this research gap by studying the influence of XAI on the trust of software developers in AI-assisted code reviews. Method: We conducted a within-subjects user study with 34 participants from diverse programming backgrounds, comparing three experimental LLM-based automated code review systems with varying levels of XAI support: Condition A (detailed explanation and review feedback), Condition B (review feedback only), and Condition C (no explanations). Participants were shown a series of real-world code change requests along with the AI-generated code reviews. During the study, we measured trust perceptions for each system using a questionnaire, agreement with the AI recommendation, the reasoning for accepting or rejecting the code change, and the time taken to review the code change. Results: Our quantitative results show that the level of explanation significantly influences both the level of trust of software developers and their agreement with AI recommendations, but in different ways. Full explanations (Condition A) yield the highest perceived trust (M = 3.99/5) but not the highest agreement with AI recommendations, whereas moderate explanations (Condition B) achieve the highest agreement with AI (89.22%). This could suggest that more explanations prompt developers to question AI recommendations more frequently. In contrast, providing no explanations (Condition C) results in the lowest levels of trust and agreement. We also find that the level of explanation did not significantly impact the time taken to accept or reject a code change. Across all conditions, the most commonly cited reasons for code change decisions were changes in code readability and the correctness of the implementation. Conclusion: Overall, these findings indicate that incorporating XAI into the code review process significantly changes the trust perceptions and agreement with AI recommendations for software developers. These results provide insights for the design and evaluation of trustworthy AI-based code review systems, and support researchers in the design of studies on the human factors of AI-assisted software development.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 697d23b6-dda6-4d01-86d3-001834b4d88aBuilds on2
- Trust in Collaborative Automation in High Stakes Software Engineering Work: A Case Study at NASADavid Gray Widder, Laura Dabbish, James D. Herbsleb, Alexandra Holloway et al.CHI 2021 · 18 citations
- EvaCRC: Evaluating Code Review CommentsLanxin Yang, Jinwei Xu, Yifan Zhang, He Zhang et al.FSE 2023 · 14 citations
Related papers
- Deep Learning-based Code Reviews: A Paradigm Shift or a Double-Edged Sword?Rosalia Tufano, Alberto Martin-Lopez, Ahmad Tayeb, Ozren Dabic et al.ICSE 2025 · 1 citation
- TACO: Trust Assessment of Large Language Models in Coding Assistance TasksShihao Weng, Yang Feng, Jincheng Li, Yining Yin et al.ICSE 2026
- "Help Me Help the AI": Understanding How Explainability Can Support Human-AI InteractionSunnie S. Y. Kim, Elizabeth Anne Watkins, Olga Russakovsky, Ruth Fong et al.CHI 2023 · 178 citations
- Trust Dynamics in AI-Assisted Development: Definitions, Factors, and ImplicationsSadra Sabouri, Philipp Eibl, Xinyi Zhou, Morteza Ziyadi et al.ICSE 2025 · 2 citations
- Enabling Global, Human-Centered Explanations for LLMs: From Tokens to Interpretable Code and Test GenerationDipin Khati, Daniel Rodriguez-Cardenas, David N. Palacio, Alejandro Velasco et al.ICSE 2026
