Do Input Gradients Highlight Discriminative Features?
Harshay Shah, Prateek Jain, Praneeth Netrapalli
摘要
Post-hoc gradient-based interpretability methods [1, 2] that provide instancespecific explanations of model predictions are often based on assumption (A): magnitude of input gradients-gradients of logits with respect to input-noisily highlight discriminative task-relevant features. In this work, we test the validity of assumption (A) using a three-pronged approach: 1. We develop an evaluation framework, DiffROAR, to test assumption (A) on four image classification benchmarks. Our results suggest that (i) input gradients of standard models (i.e., trained on original data) may grossly violate (A), whereas (ii) input gradients of adversarially robust models satisfy (A) reasonably well. 2. We then introduce BlockMNIST, an MNIST-based semi-real dataset, that by design encodes a priori knowledge of discriminative features. Our analysis on BlockMNIST leverages this information to validate as well as characterize differences between input gradient attributions of standard and robust models. 3. Finally, we theoretically prove that our empirical findings hold on a simplified version of the BlockMNIST dataset. Specifically, we prove that input gradients of standard one-hidden-layer MLPs trained on this dataset do not highlight instance-specific "signal" coordinates, thus grossly violating (A). Our findings motivate the need to formalize and test common assumptions in interpretability in a falsifiable manner [3] . We believe that the DiffROAR framework and BlockMNIST datasets serve as sanity checks to audit interpretability methods; code and data available at https://github.com/harshays/inputgradients . * Part of the work completed after joining Google Research India 2 In Appendix C, we show that our results also hold for input gradients taken w.r.t. the loss 35th Conference on Neural Information Processing Systems (NeurIPS 2021).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper26
- Protein Design with Guided Discrete DiffusionNate Gruver, Samuel Stanton, Nathan C. Frey, Tim G. J. Rudner 等NeurIPS 2023 · 被引用 246 次
- A Consistent and Efficient Evaluation Strategy for Attribution MethodsYao Rong, Tobias Leemann, Vadim Borisov, Gjergji Kasneci 等ICML 2022 · 被引用 138 次
- B-cos Networks: Alignment is All We Need for InterpretabilityMoritz Böhle, Mario Fritz, Bernt SchieleCVPR 2022 · 被引用 62 次
- ModelDiff: A Framework for Comparing Learning AlgorithmsHarshay Shah, Sung Min Park, Andrew Ilyas, Aleksander MadryICML 2023 · 被引用 36 次
- Towards Better Understanding Attribution MethodsSukrut Rao, Moritz Böhle, Bernt SchieleCVPR 2022 · 被引用 32 次
它引用的顶会 Paper11
- On Adaptive Attacks to Adversarial Example DefensesFlorian Tramèr, Nicholas Carlini, Wieland Brendel, Aleksander MadryNeurIPS 2020 · 被引用 1,026 次
- Manipulating and Measuring Model InterpretabilityForough Poursabzi-Sangdeh, Daniel G. Goldstein, Jake M. Hofman, Jennifer Wortman Vaughan 等CHI 2021 · 被引用 663 次
- Do Adversarially Robust ImageNet Models Transfer Better?Hadi Salman, Andrew Ilyas, Logan Engstrom, Ashish Kapoor 等NeurIPS 2020 · 被引用 506 次
- The Pitfalls of Simplicity Bias in Neural NetworksHarshay Shah, Kaustav Tamuly, Aditi Raghunathan, Prateek Jain 等NeurIPS 2020 · 被引用 503 次
- Debugging Tests for Model ExplanationsJulius Adebayo, Michael Muelly, Ilaria Liccardi, Been KimNeurIPS 2020 · 被引用 209 次
相关 Paper
- Pixel-level Certified Explanations via Randomized SmoothingAlaa Anani, Tobias Lorenz, Mario Fritz, Bernt SchieleICML 2025
- Interpreting Robustness Proofs of Deep Neural NetworksDebangshu Banerjee, Avaljot Singh, Gagandeep SinghICLR 2024 · 被引用 6 次
- Robust Models Are More Interpretable Because Attributions Look NormalZifan Wang, Matt Fredrikson, Anupam DattaICML 2022 · 被引用 33 次
- Post hoc Explanations may be Ineffective for Detecting Unknown Spurious CorrelationJulius Adebayo, Michael Muelly, Harold Abelson, Been KimICLR 2022 · 被引用 102 次
- Overinterpretation reveals image classification model pathologiesBrandon Carter, Siddhartha Jain, Jonas Mueller, David GiffordNeurIPS 2021 · 被引用 59 次
