Incorporating Attribution Importance for Improving Faithfulness Metrics
Zhixue Zhao, Nikolaos Aletras
摘要
Feature attribution methods (FAs) are popular approaches for providing insights into the model reasoning process of making predictions. The more faithful a FA is, the more accurately it reflects which parts of the input are more important for the prediction. Widely used faithfulness metrics, such as sufficiency and comprehensiveness use a hard erasure criterion, i.e. entirely removing or retaining the top most important tokens ranked by a given FA and observing the changes in predictive likelihood. However, this hard criterion ignores the importance of each individual token, treating them all equally for computing sufficiency and comprehensiveness. In this paper, we propose a simple yet effective soft erasure criterion. Instead of entirely removing or retaining tokens from the input, we randomly mask parts of the token vector representations proportionately to their FA importance. Extensive experiments across various natural language processing tasks and different FAs show that our soft-sufficiency and softcomprehensiveness metrics consistently prefer more faithful explanations compared to hard sufficiency and comprehensiveness. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Why and How LLMs Hallucinate: Connecting the Dots with Subsequence AssociationsYiyou Sun, Yu Gai, Lijie Chen, Abhilasha Ravichander 等NeurIPS 2025 · 被引用 20 次
- Hessian-Enhanced Token Attribution (HETA): Interpreting Autoregressive LLMsVishal Pramanik, Maisha Maliha, Nathaniel D. Bastian, Sumit Kumar JhaICLR 2026 · 被引用 3 次
- RULEBREAKERS: Challenging LLMs at the Crossroads between Formal Logic and Human-like ReasoningJason Chan, Robert J. Gaizauskas, Zhixue ZhaoICML 2025
它引用的顶会 Paper9
- A Diagnostic Study of Explainability Techniques for Text ClassificationPepa Atanasova, Jakob Grue Simonsen, Christina Lioma, Isabelle AugensteinEMNLP 2020 · 被引用 158 次
- Evaluating and Characterizing Human RationalesSamuel Carton, Anirudh Rathore, Chenhao TanEMNLP 2020 · 被引用 38 次
- ERASER: A Benchmark to Evaluate Rationalized NLP ModelsJay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric P. Lehman 等ACL 2020 · 被引用 36 次
- On the Sensitivity and Stability of Model Interpretations in NLPFan Yin, Zhouxing Shi, Cho-Jui Hsieh, Kai-Wei ChangACL 2022 · 被引用 35 次
- Connecting Attributions and QA Model Behavior on Realistic CounterfactualsXi Ye, Rohan Nair, Greg DurrettEMNLP 2021 · 被引用 13 次
相关 Paper
- Faithfulness Measurable Masked Language ModelsAndreas Madsen, Siva Reddy, Sarath ChandarICML 2024 · 被引用 6 次
- A Comparative Study of Faithfulness Metrics for Model Interpretability MethodsChun Sik Chan, Huanqi Kong, Guanqing LiangACL 2022
- Flexible Instance-Specific Rationalization of NLP ModelsGeorge Chrysostomou, Nikolaos AletrasAAAI 2022 · 被引用 17 次
- Towards Long-Horizon Interpretability: Efficient and Faithful Multi-Token Attribution for Reasoning LLMsWenbo Pan, Zhichao Liu, Xianlong Wang, Yu Haining 等ICML 2026 · 被引用 3 次
- The Out-of-Distribution Problem in Explainability and Search Methods for Feature Importance ExplanationsPeter Hase, Harry Xie, Mohit BansalNeurIPS 2021 · 被引用 121 次
