Unraveling the Dilemma of AI Errors: Exploring the Effectiveness of Human and Machine Explanations for Large Language Models
Marvin Pafla, Kate Larson, Mark Hancock
Abstract
The field of eXplainable artificial intelligence (XAI) has produced a plethora of methods (e.g., saliency-maps) to gain insight into artificial intelligence (AI) models, and has exploded with the rise of deep learning (DL). However, human-participant studies question the efficacy of these methods, particularly when the AI output is wrong. In this study, we collected and analyzed 156 human-generated text and saliency-based explanations collected in a question-answering task (N = 40) and compared them empirically to state-of-the-art XAI explanations (integrated gradients, conservative LRP, and ChatGPT) in a human-participant study (N = 136). Our findings show that participants found human saliency maps to be more helpful in explaining AI answers than machine saliency maps, but performance negatively correlated with trust in the AI model and explanations. This finding hints at the dilemma of AI errors in explanation, where helpful explanations can lead to lower task performance when they support wrong AI predictions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Fostering Appropriate Reliance on Large Language Models: The Role of Explanations, Sources, and InconsistenciesSunnie S. Y. Kim, Jennifer Wortman Vaughan, Q. Vera Liao, Tania Lombrozo et al.CHI 2025 · 118 citations
- How CO2STLY Is CHI? The Carbon Footprint of Generative AI in HCI Research and What We Should Do About ItNanna Inie, Jeanette Falk, Raghavendra SelvanCHI 2025 · 33 citations
- How Do HCI Researchers Study Cognitive Biases? A Scoping ReviewNattapat Boonprakong, Benjamin Tag, Jorge Gonçalves, Tilman DinglerCHI 2025 · 19 citations
- Designing Effective AI Explanations for Misinformation Detection: A Comparative Study of Content, Social, and Combined ExplanationsYeaeun Gong, Yifan Liu, Lanyu Shang, Na Wei et al.CSCW 2025 · 2 citations
Builds on13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-assisted Decision-makingZana Buçinca, Maja Barbara Malaya, Krzysztof Z. GajosCSCW 2021 · 962 citations
- Does the Whole Exceed its Parts? The Effect of AI Explanations on Complementary Team PerformanceGagan Bansal, Tongshuang Wu, Joyce Zhou, Raymond Fok et al.CHI 2021 · 713 citations
- Manipulating and Measuring Model InterpretabilityForough Poursabzi-Sangdeh, Daniel G. Goldstein, Jake M. Hofman, Jennifer Wortman Vaughan et al.CHI 2021 · 663 citations
- Re-examining Whether, Why, and How Human-AI Interaction Is Uniquely Difficult to DesignQian Yang, Aaron Steinfeld, Carolyn P. Rosé, John ZimmermanCHI 2020 · 604 citations
Related papers
- Fool Me Once? Contrasting Textual and Visual Explanations in a Clinical Decision-Support SettingMaxime Kayser, Bayar Menzat, Cornelius Emde, Bogdan Bercean et al.EMNLP 2024 · 8 citations
- A Psychological Theory of ExplainabilityScott Cheng-Hsin Yang, Tomas Folke, Patrick ShaftoICML 2022 · 21 citations
- Understanding the Effect of Counterfactual Explanations on Trust and Reliance on AI for Human-AI Collaborative Clinical Decision MakingMin Hun Lee, Chong Jun ChewCSCW 2023 · 74 citations
- Debiased-CAM to mitigate image perturbations with faithful visual explanations of machine learningWencan Zhang, Mariella Dimiccoli, Brian Y. LimCHI 2022 · 12 citations
- The Impact of Imperfect XAI on Human-AI Decision-MakingKatelyn Morrison, Philipp Spitzer, Violet Turri, Michelle Feng et al.CSCW 2024 · 60 citations
