Rethinking the Role of Gradient-based Attribution Methods for Model Interpretability
Suraj Srinivas, François Fleuret
摘要
Current methods for the interpretability of discriminative deep neural networks commonly rely on the model's input-gradients, i.e., the gradients of the output logits w.r.t. the inputs. The common assumption is that these input-gradients contain information regarding p θ (y | x), the model's discriminative capabilities, thus justifying their use for interpretability. However, in this work we show that these input-gradients can be arbitrarily manipulated as a consequence of the shiftinvariance of softmax without changing the discriminative function. This leaves an open question: if input-gradients can be arbitrary, why are they highly structured and explanatory in standard models? We investigate this by re-interpreting the logits of standard softmax-based classifiers as unnormalized log-densities of the data distribution and show that input-gradients can be viewed as gradients of a class-conditional density model p θ (x | y) implicit within the discriminative model. This leads us to hypothesize that the highly structured and explanatory nature of input-gradients may be due to the alignment of this class-conditional model p θ (x | y) with that of the ground truth data distribution p data (x | y). We test this hypothesis by studying the effect of density alignment on gradient explanations. To achieve this density alignment, we use an algorithm called score-matching, and propose novel approximations to this algorithm to enable training large-scale models. Our experiments show that improving the alignment of the implicit density model with the data distribution enhances gradient structure and explanatory power while reducing this alignment has the opposite effect. This also leads us to conjecture that unintended density alignment in standard neural network training may explain the highly structured nature of input-gradients observed in practice. Overall, our finding that input-gradients capture information regarding an implicit generative model implies that we need to re-think their use for interpreting discriminative models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Which Explanation Should I Choose? A Function Approximation Perspective to Characterizing Post Hoc ExplanationsTessa Han, Suraj Srinivas, Himabindu LakkarajuNeurIPS 2022 · 被引用 126 次
- Post hoc Explanations may be Ineffective for Detecting Unknown Spurious CorrelationJulius Adebayo, Michael Muelly, Harold Abelson, Been KimICLR 2022 · 被引用 102 次
- EDGE: Explaining Deep Reinforcement Learning PoliciesWenbo Guo, Xian Wu, Usmann Khan, Xinyu XingNeurIPS 2021 · 被引用 79 次
- B-cos Networks: Alignment is All We Need for InterpretabilityMoritz Böhle, Mario Fritz, Bernt SchieleCVPR 2022 · 被引用 62 次
- Rethinking Attention-Model Explainability through Faithfulness Violation TestYibing Liu, Haoliang Li, Yangyang Guo, Chenqi Kong 等ICML 2022 · 被引用 60 次
它引用的顶会 Paper4
- Your classifier is secretly an energy based model and you should treat it like oneWill Grathwohl, Kuan-Chieh Wang, Jörn-Henrik Jacobsen, David Duvenaud 等ICLR 2020 · 被引用 643 次
- Understanding Deep Networks via Extremal Perturbations and Smooth MasksRuth Fong, Mandela Patrick, Andrea VedaldiICCV 2019 · 被引用 480 次
- Efficient Learning of Generative Models via Finite-Difference Score MatchingTianyu Pang, Taufik Xu, Chongxuan Li, Yang Song 等NeurIPS 2020 · 被引用 67 次
- Interpretable Deep Learning under FireXinyang Zhang, Ningfei Wang, Hua Shen, Shouling Ji 等USENIX Security 2020
相关 Paper
- Efficient Score Matching with Deep Equilibrium LayersYuhao Huang, Qingsong Wang, Akwum Onwunta, Bao WangICLR 2024 · 被引用 4 次
- "Why Not Other Classes?": Towards Class-Contrastive Back-Propagation ExplanationsYipei Wang, Xiaoqian WangNeurIPS 2022 · 被引用 17 次
- Score-based generative models break the curse of dimensionality in learning a family of sub-Gaussian distributionsFrank Cole, Yulong LuICLR 2024 · 被引用 9 次
- FAIRER: Fairness as Decision Rationale AlignmentTianlin Li, Qing Guo, Aishan Liu, Mengnan Du 等ICML 2023 · 被引用 20 次
- On the explainable properties of 1-Lipschitz Neural Networks: An Optimal Transport PerspectiveMathieu Serrurier, Franck Mamalet, Thomas Fel, Louis Béthune 等NeurIPS 2023 · 被引用 11 次
