Incorporating Attribution Importance for Improving Faithfulness Metrics
Zhixue Zhao, Nikolaos Aletras
Abstract
Feature attribution methods (FAs) are popular approaches for providing insights into the model reasoning process of making predictions. The more faithful a FA is, the more accurately it reflects which parts of the input are more important for the prediction. Widely used faithfulness metrics, such as sufficiency and comprehensiveness use a hard erasure criterion, i.e. entirely removing or retaining the top most important tokens ranked by a given FA and observing the changes in predictive likelihood. However, this hard criterion ignores the importance of each individual token, treating them all equally for computing sufficiency and comprehensiveness. In this paper, we propose a simple yet effective soft erasure criterion. Instead of entirely removing or retaining tokens from the input, we randomly mask parts of the token vector representations proportionately to their FA importance. Extensive experiments across various natural language processing tasks and different FAs show that our soft-sufficiency and softcomprehensiveness metrics consistently prefer more faithful explanations compared to hard sufficiency and comprehensiveness. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d738d214-b3c4-4695-a673-be1602fe9b88Cited by top-tier papers3
- Why and How LLMs Hallucinate: Connecting the Dots with Subsequence AssociationsYiyou Sun, Yu Gai, Lijie Chen, Abhilasha Ravichander et al.NeurIPS 2025 · 20 citations
- Hessian-Enhanced Token Attribution (HETA): Interpreting Autoregressive LLMsVishal Pramanik, Maisha Maliha, Nathaniel D. Bastian, Sumit Kumar JhaICLR 2026 · 3 citations
- RULEBREAKERS: Challenging LLMs at the Crossroads between Formal Logic and Human-like ReasoningJason Chan, Robert J. Gaizauskas, Zhixue ZhaoICML 2025
Builds on9
- A Diagnostic Study of Explainability Techniques for Text ClassificationPepa Atanasova, Jakob Grue Simonsen, Christina Lioma, Isabelle AugensteinEMNLP 2020 · 158 citations
- Evaluating and Characterizing Human RationalesSamuel Carton, Anirudh Rathore, Chenhao TanEMNLP 2020 · 38 citations
- ERASER: A Benchmark to Evaluate Rationalized NLP ModelsJay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric P. Lehman et al.ACL 2020 · 36 citations
- On the Sensitivity and Stability of Model Interpretations in NLPFan Yin, Zhouxing Shi, Cho-Jui Hsieh, Kai-Wei ChangACL 2022 · 35 citations
- Connecting Attributions and QA Model Behavior on Realistic CounterfactualsXi Ye, Rohan Nair, Greg DurrettEMNLP 2021 · 13 citations
Related papers
- Faithfulness Measurable Masked Language ModelsAndreas Madsen, Siva Reddy, Sarath ChandarICML 2024 · 6 citations
- A Comparative Study of Faithfulness Metrics for Model Interpretability MethodsChun Sik Chan, Huanqi Kong, Guanqing LiangACL 2022
- Flexible Instance-Specific Rationalization of NLP ModelsGeorge Chrysostomou, Nikolaos AletrasAAAI 2022 · 17 citations
- Towards Long-Horizon Interpretability: Efficient and Faithful Multi-Token Attribution for Reasoning LLMsWenbo Pan, Zhichao Liu, Xianlong Wang, Yu Haining et al.ICML 2026 · 3 citations
- The Out-of-Distribution Problem in Explainability and Search Methods for Feature Importance ExplanationsPeter Hase, Harry Xie, Mohit BansalNeurIPS 2021 · 121 citations
