Normalized AOPC: Fixing Misleading Faithfulness Metrics for Feature Attributions Explainability
Joakim Edin, Andreas Geert Motzfeldt, Casper L. Christensen, Tuukka Ruotsalo, Lars Maaløe, Maria Maistro
摘要
Deep neural network predictions are notoriously difficult to interpret. Feature attribution methods aim to explain these predictions by identifying the contribution of each input feature. Faithfulness, often evaluated using the area over the perturbation curve (AOPC), reflects feature attributions' accuracy in describing the internal mechanisms of deep neural networks. However, many studies rely on AOPC to compare faithfulness across different models, which we show can lead to false conclusions about models' faithfulness. Specifically, we find that AOPC is sensitive to variations in the model, resulting in unreliable cross-model comparisons. Moreover, AOPC scores are difficult to interpret in isolation without knowing the model-specific lower and upper limits. To address these issues, we propose a normalization approach, Normalized AOPC (NAOPC), enabling consistent cross-model evaluations and more meaningful interpretation of individual scores. Our experiments demonstrate that this normalization can radically change AOPC results, questioning the conclusions of earlier studies and offering a more robust framework for assessing feature attribution faithfulness. Our code is available at https://github. com/JoakimEdin/naopc .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- The Out-of-Distribution Problem in Explainability and Search Methods for Feature Importance ExplanationsPeter Hase, Harry Xie, Mohit BansalNeurIPS 2021 · 被引用 121 次
- How does This Interaction Affect Me? Interpretable Attribution for Feature InteractionsMichael Tsang, Sirisha Rambhatla, Yan LiuNeurIPS 2020 · 被引用 109 次
- Learning Variational Word Masks to Improve the Interpretability of Neural Text ClassifiersHanjie Chen, Yangfeng JiEMNLP 2020 · 被引用 45 次
- ERASER: A Benchmark to Evaluate Rationalized NLP ModelsJay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric P. Lehman 等ACL 2020 · 被引用 36 次
- Discriminative Feature Attributions: Bridging Post Hoc Explainability and Inherent InterpretabilityUsha Bhalla, Suraj Srinivas, Himabindu LakkarajuNeurIPS 2023 · 被引用 18 次
相关 Paper
- Towards Better Understanding Attribution MethodsSukrut Rao, Moritz Böhle, Bernt SchieleCVPR 2022 · 被引用 32 次
- Logic Traps in Evaluating Attribution ScoresYiming Ju, Yuanzhe Zhang, Zhao Yang, Zhongtao Jiang 等ACL 2022
- MFABA: A More Faithful and Accelerated Boundary-Based Attribution Method for Deep Neural NetworksZhiyu Zhu, Huaming Chen, Jiayu Zhang, Xinyi Wang 等AAAI 2024 · 被引用 16 次
- Rethinking Attention-Model Explainability through Faithfulness Violation TestYibing Liu, Haoliang Li, Yangyang Guo, Chenqi Kong 等ICML 2022 · 被引用 60 次
- Re-calibrating Feature Attributions for Model InterpretationPeiyu Yang, Naveed Akhtar, Zeyi Wen, Mubarak Shah 等ICLR 2023
