On the Sensitivity and Stability of Model Interpretations in NLP
Fan Yin, Zhouxing Shi, Cho-Jui Hsieh, Kai-Wei Chang
摘要
Recent years have witnessed the emergence of a variety of post-hoc interpretations that aim to uncover how natural language processing (NLP) models make predictions. Despite the surge of new interpretation methods, it remains an open problem how to define and quantitatively measure the faithfulness of interpretations, i.e., to what extent interpretations reflect the reasoning process by a model. We propose two new criteria, sensitivity and stability, that provide complementary notions of faithfulness to the existed removal-based criteria. Our results show that the conclusion for how faithful interpretations are could vary substantially based on different notions. Motivated by the desiderata of sensitivity and stability, we introduce a new class of interpretation methods that adopt techniques from adversarial robustness. Empirical results show that our proposed methods are effective under the new criteria and overcome limitations of gradient-based methods on removal-based criteria. Besides text classification, we also apply interpretation methods and metrics to dependency parsing. Our results shed light on understanding the diverse set of interpretations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Faithful Explanations of Black-box NLP Models Using LLM-generated CounterfactualsYair Ori Gat, Nitay Calderon, Amir Feder, Alexander Chapanin 等ICLR 2024 · 被引用 55 次
- Faithful Vision-Language Interpretation via Concept Bottleneck ModelsSongning Lai, Lijie Hu, Junxiao Wang, Laure Berti-Équille 等ICLR 2024 · 被引用 42 次
- SEAT: Stable and Explainable AttentionLijie Hu, Yixin Liu, Ninghao Liu, Mengdi Huai 等AAAI 2023 · 被引用 30 次
- "Will You Find These Shortcuts?" A Protocol for Evaluating the Faithfulness of Input Salience Methods for Text ClassificationJasmijn Bastings, Sebastian Ebert, Polina Zablotskaia, Anders Sandholm 等EMNLP 2022 · 被引用 29 次
- Prompt-Agnostic Adversarial Perturbation for Customized Diffusion ModelsCong Wan, Yuhang He, Xiang Song, Yihong GongNeurIPS 2024 · 被引用 22 次
它引用的顶会 Paper6
- Automatic Perturbation Analysis for Scalable Certified Robustness and BeyondKaidi Xu, Zhouxing Shi, Huan Zhang, Yihan Wang 等NeurIPS 2020 · 被引用 415 次
- Evaluating Explainable AI: Which Algorithmic Explanations Help Users Predict Model Behavior?Peter Hase, Mohit BansalACL 2020 · 被引用 216 次
- A Diagnostic Study of Explainability Techniques for Text ClassificationPepa Atanasova, Jakob Grue Simonsen, Christina Lioma, Isabelle AugensteinEMNLP 2020 · 被引用 158 次
- Robustness Verification for TransformersZhouxing Shi, Huan Zhang, Kai-Wei Chang, Minlie Huang 等ICLR 2020 · 被引用 131 次
- Evaluations and Methods for Explanation through Robustness AnalysisCheng-Yu Hsieh, Chih-Kuan Yeh, Xuanqing Liu, Pradeep Kumar Ravikumar 等ICLR 2021 · 被引用 68 次
相关 Paper
- A Comparative Study of Faithfulness Metrics for Model Interpretability MethodsChun Sik Chan, Huanqi Kong, Guanqing LiangACL 2022
- A Causal Lens for Evaluating Faithfulness MetricsKerem Zaman, Shashank SrivastavaEMNLP 2025
- Probing Classifiers are Unreliable for Concept Removal and DetectionAbhinav Kumar, Chenhao Tan, Amit SharmaNeurIPS 2022 · 被引用 46 次
- Towards Rigorous Interpretations: a Formalisation of Feature AttributionDarius Afchar, Vincent Guigue, Romain HennequinICML 2021 · 被引用 22 次
- Measuring Association Between Labels and Free-Text RationalesSarah Wiegreffe, Ana Marasovic, Noah A. SmithEMNLP 2021 · 被引用 12 次
