Evaluating the Impact of Explainable AI on Trust in AI-Assisted Code Review
Zhenhan Gao, Marvin Muñoz Barón, Umm-e Habiba, Daniel Graziotin, Stefan Wagner
摘要
Background: The use of large language models (LLMs) for automated code review has brought significant change to a time-consuming part of software engineering. Prior work has shown that LLM-based code tools can improve code quality and enable more robust software development processes. As the tools get more powerful, the explanations behind their decisions remain hard to understand. Developers struggle to assess the validity of LLM-generated code reviews, making it difficult to gauge how much trust they should place in them. While the application of automated code review with LLMs has been extensively investigated, the inclusion of Explainable AI (XAI) for transparency in code reviews and its impact on trust are yet to be explored. Objective: We aim to address this research gap by studying the influence of XAI on the trust of software developers in AI-assisted code reviews. Method: We conducted a within-subjects user study with 34 participants from diverse programming backgrounds, comparing three experimental LLM-based automated code review systems with varying levels of XAI support: Condition A (detailed explanation and review feedback), Condition B (review feedback only), and Condition C (no explanations). Participants were shown a series of real-world code change requests along with the AI-generated code reviews. During the study, we measured trust perceptions for each system using a questionnaire, agreement with the AI recommendation, the reasoning for accepting or rejecting the code change, and the time taken to review the code change. Results: Our quantitative results show that the level of explanation significantly influences both the level of trust of software developers and their agreement with AI recommendations, but in different ways. Full explanations (Condition A) yield the highest perceived trust (M = 3.99/5) but not the highest agreement with AI recommendations, whereas moderate explanations (Condition B) achieve the highest agreement with AI (89.22%). This could suggest that more explanations prompt developers to question AI recommendations more frequently. In contrast, providing no explanations (Condition C) results in the lowest levels of trust and agreement. We also find that the level of explanation did not significantly impact the time taken to accept or reject a code change. Across all conditions, the most commonly cited reasons for code change decisions were changes in code readability and the correctness of the implementation. Conclusion: Overall, these findings indicate that incorporating XAI into the code review process significantly changes the trust perceptions and agreement with AI recommendations for software developers. These results provide insights for the design and evaluation of trustworthy AI-based code review systems, and support researchers in the design of studies on the human factors of AI-assisted software development.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper2
相关 Paper
- Deep Learning-based Code Reviews: A Paradigm Shift or a Double-Edged Sword?Rosalia Tufano, Alberto Martin-Lopez, Ahmad Tayeb, Ozren Dabic 等ICSE 2025 · 被引用 1 次
- TACO: Trust Assessment of Large Language Models in Coding Assistance TasksShihao Weng, Yang Feng, Jincheng Li, Yining Yin 等ICSE 2026
- "Help Me Help the AI": Understanding How Explainability Can Support Human-AI InteractionSunnie S. Y. Kim, Elizabeth Anne Watkins, Olga Russakovsky, Ruth Fong 等CHI 2023 · 被引用 178 次
- Trust Dynamics in AI-Assisted Development: Definitions, Factors, and ImplicationsSadra Sabouri, Philipp Eibl, Xinyi Zhou, Morteza Ziyadi 等ICSE 2025 · 被引用 2 次
- Enabling Global, Human-Centered Explanations for LLMs: From Tokens to Interpretable Code and Test GenerationDipin Khati, Daniel Rodriguez-Cardenas, David N. Palacio, Alejandro Velasco 等ICSE 2026
