Improving Deep Learning Interpretability by Saliency Guided Training
Aya Abdelsalam Ismail, Héctor Corrada Bravo, Soheil Feizi
Abstract
Saliency methods have been widely used to highlight important input features in model predictions. Most existing methods use backpropagation on a modified gradient function to generate saliency maps. Thus, noisy gradients can result in unfaithful feature attributions. In this paper, we tackle this issue and introduce a saliency guided trainingprocedure for neural networks to reduce noisy gradients used in predictions while retaining the predictive performance of the model. Our saliency guided training procedure iteratively masks features with small and potentially noisy gradients while maximizing the similarity of model outputs for both masked and unmasked inputs. We apply the saliency guided training procedure to various synthetic and real data sets from computer vision, natural language processing, and time series across diverse neural architectures, including Recurrent Neural Networks, Convolutional Networks, and Transformers. Through qualitative and quantitative evaluations, we show that saliency guided training procedure significantly improves model interpretability across various domains while preserving its predictive performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8dc018ec-0692-4943-b417-9aa573be9501Cited by top-tier papers21
- Look where you look! Saliency-guided Q-networks for generalization in visual Reinforcement LearningDavid Bertoin, Adil Zouitine, Mehdi Zouitine, Emmanuel RachelsonNeurIPS 2022 · 67 citations
- Encoding Time-Series Explanations through Self-Supervised Model Behavior ConsistencyOwen Queen, Tom Hartvigsen, Teddy Koker, Huan He et al.NeurIPS 2023 · 55 citations
- UNIREX: A Unified Learning Framework for Language Model Rationale ExtractionAaron Chan, Maziar Sanjabi, Lambert Mathias, Liang Tan et al.ICML 2022 · 48 citations
- TimeX++: Learning Time-Series Explanations with Information BottleneckZichuan Liu, Tianchun Wang, Jimeng Shi, Xu Zheng et al.ICML 2024 · 33 citations
- VisFIS: Visual Feature Importance Supervision with Right-for-the-Right-Reason ObjectivesZhuofan Ying, Peter Hase, Mohit BansalNeurIPS 2022 · 16 citations
Builds on5
- Benchmarking Deep Learning Interpretability in Time Series PredictionsAya Abdelsalam Ismail, Mohamed K. Gunady, Héctor Corrada Bravo, Soheil FeiziNeurIPS 2020 · 249 citations
- Debugging Tests for Model ExplanationsJulius Adebayo, Michael Muelly, Ilaria Liccardi, Been KimNeurIPS 2020 · 209 citations
- Sanity Checks for Saliency MetricsRichard Tomsett, Dan Harborne, Supriyo Chakraborty, Prudhvi Gurram et al.AAAI 2020 · 204 citations
- Sharpen Focus: Learning With Attention Separability and ConsistencyLezi Wang, Ziyan Wu, Srikrishna Karanam, Kuan-Chuan Peng et al.ICCV 2019 · 37 citations
- ERASER: A Benchmark to Evaluate Rationalized NLP ModelsJay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric P. Lehman et al.ACL 2020 · 36 citations
Related papers
- Time series saliency maps: Explaining models across multiple domainsChristodoulos Kechris, Jonathan Dan, David AtienzaICML 2026 · 6 citations
- DANCE: Enhancing saliency maps using decoysYang Young Lu, Wenbo Guo, Xinyu Xing, William Stafford NobleICML 2021 · 14 citations
- Structured Gradient-Based Interpretations via Norm-Regularized Adversarial TrainingShizhan Gong, Qi Dou, Farzan FarniaCVPR 2024
- Saliency is a Possible Red Herring When Diagnosing Poor GeneralizationJoseph D. Viviano, Becks Simpson, Francis Dutil, Yoshua Bengio et al.ICLR 2021 · 46 citations
- Backdoor Attacks on the DNN Interpretation SystemShihong Fang, Anna ChoromanskaAAAI 2022 · 22 citations
