Framework for Evaluating Faithfulness of Local Explanations
Sanjoy Dasgupta, Nave Frost, Michal Moshkovitz
Abstract
We study the faithfulness of an explanation system to the underlying prediction model. We show that this can be captured by two properties, consistency and sufficiency, and introduce quantitative measures of the extent to which these hold. Interestingly, these measures depend on the test-time data distribution. For a variety of existing explanation systems, such as anchors, we analytically study these quantities. We also provide estimators and sample complexity bounds for empirically determining the faithfulness of black-box explanation systems. Finally, we experimentally validate the new properties and estimators.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c5e6e8f4-4f0a-4b0c-b2b0-f6ffa44ab011Cited by top-tier papers16
- Faithful Explanations of Black-box NLP Models Using LLM-generated CounterfactualsYair Ori Gat, Nitay Calderon, Amir Feder, Alexander Chapanin et al.ICLR 2024 · 55 citations
- Encoding Time-Series Explanations through Self-Supervised Model Behavior ConsistencyOwen Queen, Tom Hartvigsen, Teddy Koker, Huan He et al.NeurIPS 2023 · 55 citations
- SAFARI: Versatile and Efficient Evaluations for Robustness of InterpretabilityWei Huang, Xingyu Zhao, Gaojie Jin, Xiaowei HuangICCV 2023 · 41 citations
- X-CHAR: A Concept-based Explainable Complex Human Activity Recognition ModelJeya Vikranth Jeyakumar, Ankur Sarker, Luis Antonio Garcia, Mani B. SrivastavaUbiComp 2023 · 41 citations
- MultiViz: Towards Visualizing and Understanding Multimodal ModelsPaul Pu Liang, Yiwei Lyu, Gunjan Chhablani, Nihal Jain et al.ICLR 2023 · 15 citations
Builds on7
- Manipulating and Measuring Model InterpretabilityForough Poursabzi-Sangdeh, Daniel G. Goldstein, Jake M. Hofman, Jennifer Wortman Vaughan et al.CHI 2021 · 663 citations
- Explainable k-Means and k-Medians ClusteringMichal Moshkovitz, Sanjoy Dasgupta, Cyrus Rashtchian, Nave FrostICML 2020 · 184 citations
- Counterfactual Explanations Can Be ManipulatedDylan Slack, Anna Hilgard, Himabindu Lakkaraju, Sameer SinghNeurIPS 2021 · 182 citations
- Model Interpretability through the lens of Computational ComplexityPablo Barceló, Mikaël Monet, Jorge Pérez, Bernardo SubercaseauxNeurIPS 2020 · 135 citations
- Fairwashing explanations with off-manifold detergentChristopher J. Anders, Plamen Pasliev, Ann-Kathrin Dombrowski, Klaus-Robert Müller et al.ICML 2020 · 104 citations
Related papers
- Diagnostics-Guided Explanation GenerationPepa Atanasova, Jakob Grue Simonsen, Christina Lioma, Isabelle AugensteinAAAI 2022 · 10 citations
- On Measuring Faithfulness or Self-consistency of Natural Language ExplanationsLetitia Parcalabescu, Anette FrankACL 2024
- What LLMs Explain Is Not What They Believe: Evaluating Explanation Sufficiency Under Models' Own Input BeliefsNhi Nguyen, Shauli Ravfogel, Rajesh RanganathICML 2026
- A Necessary Step toward Faithfulness: Measuring and Improving Consistency in Free-Text ExplanationsLingjun Zhao, Hal Daumé IIIEMNLP 2025 · 3 citations
- Evaluating Explanation Methods for Neural Machine TranslationJierui Li, Lemao Liu, Huayang Li, Guanlin Li et al.ACL 2020 · 24 citations
