Framework for Evaluating Faithfulness of Local Explanations
Sanjoy Dasgupta, Nave Frost, Michal Moshkovitz
2022年份
87被引次数
16顶会引用
摘要
We study the faithfulness of an explanation system to the underlying prediction model. We show that this can be captured by two properties, consistency and sufficiency, and introduce quantitative measures of the extent to which these hold. Interestingly, these measures depend on the test-time data distribution. For a variety of existing explanation systems, such as anchors, we analytically study these quantities. We also provide estimators and sample complexity bounds for empirically determining the faithfulness of black-box explanation systems. Finally, we experimentally validate the new properties and estimators.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Faithful Explanations of Black-box NLP Models Using LLM-generated CounterfactualsYair Ori Gat, Nitay Calderon, Amir Feder, Alexander Chapanin 等ICLR 2024 · 被引用 55 次
- Encoding Time-Series Explanations through Self-Supervised Model Behavior ConsistencyOwen Queen, Tom Hartvigsen, Teddy Koker, Huan He 等NeurIPS 2023 · 被引用 55 次
- SAFARI: Versatile and Efficient Evaluations for Robustness of InterpretabilityWei Huang, Xingyu Zhao, Gaojie Jin, Xiaowei HuangICCV 2023 · 被引用 41 次
- X-CHAR: A Concept-based Explainable Complex Human Activity Recognition ModelJeya Vikranth Jeyakumar, Ankur Sarker, Luis Antonio Garcia, Mani B. SrivastavaUbiComp 2023 · 被引用 41 次
- MultiViz: Towards Visualizing and Understanding Multimodal ModelsPaul Pu Liang, Yiwei Lyu, Gunjan Chhablani, Nihal Jain 等ICLR 2023 · 被引用 15 次
它引用的顶会 Paper7
- Manipulating and Measuring Model InterpretabilityForough Poursabzi-Sangdeh, Daniel G. Goldstein, Jake M. Hofman, Jennifer Wortman Vaughan 等CHI 2021 · 被引用 663 次
- Explainable k-Means and k-Medians ClusteringMichal Moshkovitz, Sanjoy Dasgupta, Cyrus Rashtchian, Nave FrostICML 2020 · 被引用 184 次
- Counterfactual Explanations Can Be ManipulatedDylan Slack, Anna Hilgard, Himabindu Lakkaraju, Sameer SinghNeurIPS 2021 · 被引用 182 次
- Model Interpretability through the lens of Computational ComplexityPablo Barceló, Mikaël Monet, Jorge Pérez, Bernardo SubercaseauxNeurIPS 2020 · 被引用 135 次
- Fairwashing explanations with off-manifold detergentChristopher J. Anders, Plamen Pasliev, Ann-Kathrin Dombrowski, Klaus-Robert Müller 等ICML 2020 · 被引用 104 次
相关 Paper
- Diagnostics-Guided Explanation GenerationPepa Atanasova, Jakob Grue Simonsen, Christina Lioma, Isabelle AugensteinAAAI 2022 · 被引用 10 次
- On Measuring Faithfulness or Self-consistency of Natural Language ExplanationsLetitia Parcalabescu, Anette FrankACL 2024
- What LLMs Explain Is Not What They Believe: Evaluating Explanation Sufficiency Under Models' Own Input BeliefsNhi Nguyen, Shauli Ravfogel, Rajesh RanganathICML 2026
- A Necessary Step toward Faithfulness: Measuring and Improving Consistency in Free-Text ExplanationsLingjun Zhao, Hal Daumé IIIEMNLP 2025 · 被引用 3 次
- Evaluating Explanation Methods for Neural Machine TranslationJierui Li, Lemao Liu, Huayang Li, Guanlin Li 等ACL 2020 · 被引用 24 次
