DANCE: Enhancing saliency maps using decoys
Yang Young Lu, Wenbo Guo, Xinyu Xing, William Stafford Noble
Abstract
Saliency methods can make deep neural network predictions more interpretable by identifying a set of critical features in an input sample, such as pixels that contribute most strongly to a prediction made by an image classifier. Unfortunately, recent evidence suggests that many saliency methods poorly perform, especially in situations where gradients are saturated, inputs contain adversarial perturbations, or predictions rely upon inter-feature dependence. To address these issues, we propose a framework that improves the robustness of saliency methods by following a two-step procedure. First, we introduce a perturbation mechanism that subtly varies the input sample without changing its intermediate representations. Using this approach, we can gather a corpus of perturbed data samples while ensuring that the perturbed and original input samples follow the same distribution. Second, we compute saliency maps for the perturbed samples and propose a new method to aggregate saliency maps. With this design, we offset the gradient saturation influence upon interpretation. From a theoretical perspective, we show the aggregated saliency map could not only capture inter-feature dependence but, more importantly, robustify interpretation against previously described adversarial perturbation methods. Following our theoretical analysis, we present experimental results suggesting that, both qualitatively and quantitatively, our saliency method outperforms existing methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 02308f64-8c8d-4561-9609-df81c17cbab8Cited by top-tier papers6
- EDGE: Explaining Deep Reinforcement Learning PoliciesWenbo Guo, Xian Wu, Usmann Khan, Xinyu XingNeurIPS 2021 · 79 citations
- Learning to Identify Critical States for Reinforcement Learning from VideosHaozhe Liu, Mingchen Zhuge, Bing Li, Yuhui Wang et al.ICCV 2023 · 14 citations
- Training for Stable Explanation for FreeChao Chen, Chenghua Guo, Rufeng Chen, Guixiang Ma et al.NeurIPS 2024 · 7 citations
- Path Choice Matters for Clear Attributions in Path MethodsBorui Zhang, Wenzhao Zheng, Jie Zhou, Jiwen LuICLR 2024 · 5 citations
- AdaptGrad: Adaptive Sampling to Reduce NoiseLinjiang Zhou, Chao Ma, Zepeng Wang, Libing Wu et al.NeurIPS 2025 · 3 citations
Builds on3
- Restricting the Flow: Information Bottlenecks for AttributionKarl Schulz, Leon Sixt, Federico Tombari, Tim LandgrafICLR 2020 · 220 citations
- Concise Explanations of Neural Networks using Adversarial TrainingPrasad Chalasani, Jiefeng Chen, Amrita Roy Chowdhury, Xi Wu et al.ICML 2020 · 148 citations
- When Explanations Lie: Why Many Modified BP Attributions FailLeon Sixt, Maximilian Granz, Tim LandgrafICML 2020 · 147 citations
Related papers
- Attack to Explain Deep RepresentationMohammad A. A. K. Jalwana, Naveed Akhtar, Mohammed Bennamoun, Ajmal MianCVPR 2020
- CAMERAS: Enhanced Resolution and Sanity Preserving Class Activation Mapping for Image SaliencyMohammad A. A. K. Jalwana, Naveed Akhtar, Mohammed Bennamoun, Ajmal MianCVPR 2021
- One step further: evaluating interpreters using metamorphic testingMing Fan, Jiali Wei, Wuxia Jin, Zhou Xu et al.ISSTA 2022 · 7 citations
- Robust Superpixel-Guided Attentional Adversarial AttackXiaoyi Dong, Jiangfan Han, Dongdong Chen, Jiayang Liu et al.CVPR 2020
- Building Reliable Explanations of Unreliable Neural Networks: Locally Smoothing Perspective of Model InterpretationDohun Lim, Hyeonseok Lee, Sungchan KimCVPR 2021
