Explaining Model Confidence Using Counterfactuals
Thao Le, Tim Miller, Ronal Singh, Liz Sonenberg
摘要
Displaying confidence scores in human-AI interaction has been shown to help build trust between humans and AI systems. However, most existing research uses only the confidence score as a form of communication. As confidence scores are just another model output, users may want to understand why the algorithm is confident to determine whether to accept the confidence score. In this paper, we show that counterfactual explanations of confidence scores help study participants to better understand and better trust a machine learning model's prediction. We present two methods for understanding model confidence using counterfactual explanation: (1) based on counterfactual examples; and (2) based on visualisation of the counterfactual space. Both increase understanding and trust for study participants over a baseline of no explanation, but qualitative results show that they are used quite differently, leading to recommendations of when to use each one and directions of designing better explanations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Show Me the Work: Fact-Checkers' Requirements for Explainable Automated Fact-CheckingGreta Warren, Irina Shklovski, Isabelle AugensteinCHI 2025 · 被引用 15 次
- Sometimes You Need Facts, and Sometimes a Hug: Understanding Older Adults' Preferences for Explanations in LLM-Based Conversational AI SystemsNiharika Mathur, Tamara Zubatiy, Agata Rozga, Jodi Forlizzi 等CHI 2026 · 被引用 4 次
它引用的顶会 Paper3
- Does the Whole Exceed its Parts? The Effect of AI Explanations on Complementary Team PerformanceGagan Bansal, Tongshuang Wu, Joyce Zhou, Raymond Fok 等CHI 2021 · 被引用 713 次
- Getting a CLUE: A Method for Explaining Uncertainty EstimatesJavier Antorán, Umang Bhatt, Tameem Adel, Adrian Weller 等ICLR 2021 · 被引用 41 次
- Contrastive Explanations for Model InterpretabilityAlon Jacovi, Swabha Swayamdipta, Shauli Ravfogel, Yanai Elazar 等EMNLP 2021 · 被引用 12 次
相关 Paper
- Impact of Explanation Techniques and Representations on Users' Comprehension and Confidence in Explainable AIJulien Delaunay, Luis Galárraga, Christine Largouët, Niels van BerkelCSCW 2025 · 被引用 6 次
- DECE: Decision Explorer with Counterfactual Explanations for Machine Learning ModelsFurui Cheng, Yao Ming, Huamin QuIEEE VIS 2020 · 被引用 118 次
- Explainable Reinforcement Learning through a Causal LensPrashan Madumal, Tim Miller, Liz Sonenberg, Frank VetereAAAI 2020 · 被引用 408 次
- Understanding the Effect of Counterfactual Explanations on Trust and Reliance on AI for Human-AI Collaborative Clinical Decision MakingMin Hun Lee, Chong Jun ChewCSCW 2023 · 被引用 74 次
- Towards Unifying Evaluation of Counterfactual Explanations: Leveraging Large Language Models for Human-Centric AssessmentsMarharyta Domnich, Julius Välja, Rasmus Moorits Veski, Giacomo Magnifico 等AAAI 2025 · 被引用 11 次
