Concise Explanations of Neural Networks using Adversarial Training
Prasad Chalasani, Jiefeng Chen, Amrita Roy Chowdhury, Xi Wu, Somesh Jha
Abstract
We show new connections between adversarial learning and explainability for deep neural networks (DNNs). One form of explanation of the output of a neural network model in terms of its input features, is a vector of feature-attributions. Two desirable characteristics of an attributionbased explanation are: (1) sparseness: the attributions of irrelevant or weakly relevant features should be negligible, thus resulting in concise explanations in terms of the significant features, and (2) stability: it should not vary significantly within a small local neighborhood of the input. Our first contribution is a theoretical exploration of how these two properties (when using attributions based on Integrated Gradients, or IG) are related to adversarial training, for a class of 1-layer networks (which includes logistic regression models for binary and multi-class classification); for these networks we show that (a) adversarial training using an ∞ -bounded adversary produces models with sparse attribution vectors, and (b) natural model-training while encouraging stable explanations (via an extra term in the loss function), is equivalent to adversarial training. Our second contribution is an empirical verification of phenomenon (a), which we show, somewhat surprisingly, occurs not only in 1-layer networks, but also DNNs trained on standard image datasets, and extends beyond IGbased attributions, to those based on DeepSHAP: adversarial training with ∞ -bounded perturbations yields significantly sparser attribution vectors, with little degradation in performance on natural test data, compared to natural training. Moreover, the sparseness of the attribution vectors is significantly better than that achievable via 1 -regularized natural training.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers35
- Reliable Post hoc Explanations: Modeling Uncertainty in ExplainabilityDylan Slack, Anna Hilgard, Sameer Singh, Himabindu LakkarajuNeurIPS 2021 · 240 citations
- Do Input Gradients Highlight Discriminative Features?Harshay Shah, Prateek Jain, Praneeth NetrapalliNeurIPS 2021 · 74 citations
- Towards the Unification and Robustness of Perturbation and Gradient Based ExplanationsSushant Agarwal, Shahin Jabbari, Chirag Agarwal, Sohini Upadhyay et al.ICML 2021 · 71 citations
- Adversarially Robust 3D Point Cloud Recognition Using Self-SupervisionsJiachen Sun, Yulong Cao, Christopher B. Choy, Zhiding Yu et al.NeurIPS 2021 · 64 citations
- NoiseGrad - Enhancing Explanations by Introducing Stochasticity to Model WeightsKirill Bykov, Anna Hedström, Shinichi Nakajima, Marina M.-C. HöhneAAAI 2022 · 43 citations
Related papers
- Smoothed Geometry for Robust AttributionZifan Wang, Haofan Wang, Shakul Ramkumar, Piotr Mardziel et al.NeurIPS 2020 · 67 citations
- Interpreting Attributions and Interactions of Adversarial AttacksXin Wang, Shuyun Lin, Hao Zhang, Yufei Zhu et al.ICCV 2021 · 20 citations
- Guided Integrated Gradients: An Adaptive Path Method for Removing NoiseAndrei Kapishnikov, Subhashini Venugopalan, Besim Avci, Ben Wedin et al.CVPR 2021
- Structured Gradient-Based Interpretations via Norm-Regularized Adversarial TrainingShizhan Gong, Qi Dou, Farzan FarniaCVPR 2024
- Understanding Deep Networks via Extremal Perturbations and Smooth MasksRuth Fong, Mandela Patrick, Andrea VedaldiICCV 2019 · 480 citations
