A Psychological Theory of Explainability
Scott Cheng-Hsin Yang, Tomas Folke, Patrick Shafto
摘要
The goal of explainable Artificial Intelligence (XAI) is to generate human-interpretable explanations, but there are no computationally precise theories of how humans interpret AI generated explanations. The lack of theory means that validation of XAI must be done empirically, on a case-by-case basis, which prevents systematic theory-building in XAI. We propose a psychological theory of how humans draw conclusions from saliency maps, the most common form of XAI explanation, which for the first time allows for precise prediction of explainee inference conditioned on explanation. Our theory posits that absent explanation humans expect the AI to make similar decisions to themselves, and that they interpret an explanation by comparison to the explanations they themselves would give. Comparison is formalized via Shepard's universal law of generalization in a similarity space, a classic theory from cognitive science. A pre-registered user study on AI image classifications with saliency map explanations demonstrate that our theory quantitatively matches participants' predictions of the AI.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- People Attribute Purpose to Autonomous Vehicles When Explaining Their Behavior: Insights from Cognitive Science for Explainable AIBalint Gyevnar, Stephanie Droop, Tadeg Quillien, Shay B. Cohen 等CHI 2025 · 被引用 7 次
- AdaptGrad: Adaptive Sampling to Reduce NoiseLinjiang Zhou, Chao Ma, Zepeng Wang, Libing Wu 等NeurIPS 2025 · 被引用 3 次
它引用的顶会 Paper7
- Manipulating and Measuring Model InterpretabilityForough Poursabzi-Sangdeh, Daniel G. Goldstein, Jake M. Hofman, Jennifer Wortman Vaughan 等CHI 2021 · 被引用 663 次
- Interpreting Interpretability: Understanding Data Scientists' Use of Interpretability Tools for Machine LearningHarmanpreet Kaur, Harsha Nori, Samuel Jenkins, Rich Caruana 等CHI 2020 · 被引用 541 次
- Evaluating Explainable AI: Which Algorithmic Explanations Help Users Predict Model Behavior?Peter Hase, Mohit BansalACL 2020 · 被引用 216 次
- Debugging Tests for Model ExplanationsJulius Adebayo, Michael Muelly, Ilaria Liccardi, Been KimNeurIPS 2020 · 被引用 209 次
- Visualizing Deep Networks by Optimizing with Integrated GradientsZhongang Qi, Saeed Khorram, Fuxin LiAAAI 2020 · 被引用 149 次
相关 Paper
- Unraveling the Dilemma of AI Errors: Exploring the Effectiveness of Human and Machine Explanations for Large Language ModelsMarvin Pafla, Kate Larson, Mark HancockCHI 2024 · 被引用 16 次
- Fool Me Once? Contrasting Textual and Visual Explanations in a Clinical Decision-Support SettingMaxime Kayser, Bayar Menzat, Cornelius Emde, Bogdan Bercean 等EMNLP 2024 · 被引用 8 次
- Machine Semiology in Practice: Clinician Strategies for Interpreting AI-Generated Visual ExplanationsFederico Cabitza, Enrico Gallazzi, Alessia PapaleCSCW 2026
- (Mis)Communicating with our AI SystemsLaura Cros Vila, Bob L. T. SturmCHI 2025 · 被引用 2 次
- SketchXAI: A First Look at Explainability for Human SketchesZhiyu Qu, Yulia Gryaditskaya, Ke Li, Kaiyue Pang 等CVPR 2023
