Evaluating Explainable AI: Which Algorithmic Explanations Help Users Predict Model Behavior?
Peter Hase, Mohit Bansal
Abstract
Algorithmic approaches to interpreting machine learning models have proliferated in recent years. We carry out human subject tests that are the first of their kind to isolate the effect of algorithmic explanations on a key aspect of model interpretability, simulatability, while avoiding important confounding experimental factors. A model is simulatable when a person can predict its behavior on new inputs. Through two kinds of simulation tests involving text and tabular data, we evaluate five explanations methods: (1) LIME, (2) Anchor, (3) Decision Boundary, (4) a Prototype model, and (5) a Composite approach that combines explanations from each method. Clear evidence of method effectiveness is found in very few cases: LIME improves simulatability in tabular classification, and our Prototype method is effective in counterfactual simulation tests. We also collect subjective ratings of explanations, but we do not find that ratings are predictive of how helpful explanations are. Our results provide the first reliable and comprehensive estimates of how explanations influence simulatability across a variety of explanation methods and data domains. We show that (1) we need to be careful about the metrics we use to evaluate explanation methods, and (2) there is significant room for improvement in current methods. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 536b924c-7a34-4d19-a4fa-4b26a58d9e67Cited by top-tier papers64
- Does the Whole Exceed its Parts? The Effect of AI Explanations on Complementary Team PerformanceGagan Bansal, Tongshuang Wu, Joyce Zhou, Raymond Fok et al.CHI 2021 · 713 citations
- How Can I Explain This to You? An Empirical Study of Deep Neural Network Explanation MethodsJeya Vikranth Jeyakumar, Joseph Noor, Yu-Hsi Cheng, Luis Garcia et al.NeurIPS 2020 · 173 citations
- What I Cannot Predict, I Do Not Understand: A Human-Centered Evaluation Framework for Explainability MethodsJulien Colin, Thomas Fel, Rémi Cadène, Thomas SerreNeurIPS 2022 · 147 citations
- Understanding the Role of Human Intuition on Reliance in Human-AI Decision-Making with ExplanationsValerie Chen, Q. Vera Liao, Jennifer Wortman Vaughan, Gagan BansalCSCW 2023 · 146 citations
- A Consistent and Efficient Evaluation Strategy for Attribution MethodsYao Rong, Tobias Leemann, Vadim Borisov, Gjergji Kasneci et al.ICML 2022 · 138 citations
Builds on1
Related papers
- ConSim: Measuring Concept-Based Explanations' Effectiveness with Automated SimulatabilityAntonin Poché, Alon Jacovi, Agustin Martin Picard, Victor Boutin et al.ACL 2025 · 8 citations
- Use-Case-Grounded Simulations for Explanation EvaluationValerie Chen, Nari Johnson, Nicholay Topin, Gregory Plumb et al.NeurIPS 2022 · 26 citations
- Shahin: Faster Algorithms for Generating Explanations for Multiple PredictionsSona Hasani, Saravanan Thirumuruganathan, Nick Koudas, Gautam DasSIGMOD 2021 · 1 citation
- Do Models Explain Themselves? Counterfactual Simulatability of Natural Language ExplanationsYanda Chen, Ruiqi Zhong, Narutatsu Ri, Chen Zhao et al.ICML 2024 · 90 citations
- Are Human Explanations Always Helpful? Towards Objective Evaluation of Human Natural Language ExplanationsBingsheng Yao, Prithviraj Sen, Lucian Popa, James A. Hendler et al.ACL 2023 · 4 citations
