Logic Traps in Evaluating Attribution Scores
Yiming Ju, Yuanzhe Zhang, Zhao Yang, Zhongtao Jiang, Kang Liu, Jun Zhao
Abstract
Modern deep learning models are notoriously opaque, which has motivated the development of methods for interpreting how deep models predict. This goal is usually approached with attribution method, which assesses the influence of features on model predictions. As an explanation method, the evaluation criteria of attribution methods is how accurately it reflects the actual reasoning process of the model (faithfulness). Meanwhile, since the reasoning process of deep models is inaccessible, researchers design various evaluation methods to demonstrate their arguments. However, some crucial logic traps in these evaluation methods are ignored in most works, causing inaccurate evaluation and unfair comparison. This paper systematically reviews existing methods for evaluating attribution scores and summarizes the logic traps in these methods. We further conduct experiments to demonstrate the existence of each logic trap. Through both theoretical and experimental analysis, we hope to increase attention on the inaccurate evaluation of attribution scores. Moreover, with this paper, we suggest stopping focusing on improving performance under unreliable evaluation systems and starting efforts on reducing the impact of proposed logic traps.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Why and How LLMs Hallucinate: Connecting the Dots with Subsequence AssociationsYiyou Sun, Yu Gai, Lijie Chen, Abhilasha Ravichander et al.NeurIPS 2025 · 20 citations
- Explanations that reveal all through the definition of encodingAahlad Manas Puli, Nhi Nguyen, Rajesh RanganathNeurIPS 2024 · 5 citations
- Incorporating Attribution Importance for Improving Faithfulness MetricsZhixue Zhao, Nikolaos AletrasACL 2023 · 4 citations
- Normalized AOPC: Fixing Misleading Faithfulness Metrics for Feature Attributions ExplainabilityJoakim Edin, Andreas Geert Motzfeldt, Casper L. Christensen, Tuukka Ruotsalo et al.ACL 2025
- F-Fidelity: A Robust Framework for Faithfulness Evaluation of Explainable AIXu Zheng, Farhad Shirani, Zhuomin Chen, Chaohao Lin et al.ICLR 2025
Builds on9
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 1,333 citations
- Self-Attention Attribution: Interpreting Information Interactions Inside TransformerYaru Hao, Li Dong, Furu Wei, Ke XuAAAI 2021 · 282 citations
- How does This Interaction Affect Me? Interpretable Attribution for Feature InteractionsMichael Tsang, Sirisha Rambhatla, Yan LiuNeurIPS 2020 · 109 citations
- Generating Hierarchical Explanations on Text Classification via Feature Interaction DetectionHanjie Chen, Guangtao Zheng, Yangfeng JiACL 2020 · 85 citations
- Learning Variational Word Masks to Improve the Interpretability of Neural Text ClassifiersHanjie Chen, Yangfeng JiEMNLP 2020 · 45 citations
Related papers
- Towards Better Understanding Attribution MethodsSukrut Rao, Moritz Böhle, Bernt SchieleCVPR 2022 · 32 citations
- Rethinking Attention-Model Explainability through Faithfulness Violation TestYibing Liu, Haoliang Li, Yangyang Guo, Chenqi Kong et al.ICML 2022 · 60 citations
- Towards credible visual model interpretation with path attributionNaveed Akhtar, Mohammad A. A. K. JalwanaICML 2023 · 6 citations
- A Unified Taylor Framework for Revisiting Attribution MethodsHuiqi Deng, Na Zou, Mengnan Du, Weifu Chen et al.AAAI 2021 · 25 citations
- MFABA: A More Faithful and Accelerated Boundary-Based Attribution Method for Deep Neural NetworksZhiyu Zhu, Huaming Chen, Jiayu Zhang, Xinyi Wang et al.AAAI 2024 · 16 citations
