Building Reliable Explanations of Unreliable Neural Networks: Locally Smoothing Perspective of Model Interpretation
Dohun Lim, Hyeonseok Lee, Sungchan Kim
Abstract
We present a novel method for reliably explaining the predictions of neural networks. We consider an explanation reliable if it identifies input features relevant to the model output by considering the input and the neighboring data points. Our method is built on top of the assumption of smooth landscape in a loss function of the model prediction: locally consistent loss and gradient profile. A theoretical analysis established in this study suggests that those locally smooth model explanations are learned using a batch of noisy copies of the input with the L1 regularization for a saliency map. Extensive experiments support the analysis results, revealing that the proposed saliency maps retrieve the original classes of adversarial examples crafted against both naturally and adversarially trained models, significantly outperforming previous methods. We further demonstrated that such good performance results from the learning capability of this method to identify input features that are truly relevant to the model output of the input and the neighboring data points, fulfilling the requirements of a reliable explanation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ab0de513-e368-4139-a984-b519164dbbadCited by top-tier papers3
- A Causality Inspired Framework for Model InterpretationChenwang Wu, Xiting Wang, Defu Lian, Xing Xie et al.KDD 2023 · 22 citations
- MoreauGrad: Sparse and Robust Interpretation of Neural Networks via Moreau EnvelopeJingwei Zhang, Farzan FarniaICCV 2023 · 5 citations
- Improving Adversarial Robustness of Attribution via Implicit RegularizationAmir Mehrpanah, Matteo Gamba, Hossein AzizpourICML 2026
Builds on3
- XRAI: Better Attributions Through RegionsAndrei Kapishnikov, Tolga Bolukbasi, Fernanda B. Viégas, Michael TerryICCV 2019 · 251 citations
- Fooling Network Interpretation in Image ClassificationAkshayvarun Subramanya, Vipin Pillai, Hamed PirsiavashICCV 2019 · 68 citations
- There and Back Again: Revisiting Backpropagation Saliency MethodsSylvestre-Alvise Rebuffi, Ruth Fong, Xu Ji, Andrea VedaldiCVPR 2020
Related papers
- On the explainable properties of 1-Lipschitz Neural Networks: An Optimal Transport PerspectiveMathieu Serrurier, Franck Mamalet, Thomas Fel, Louis Béthune et al.NeurIPS 2023 · 11 citations
- DANCE: Enhancing saliency maps using decoysYang Young Lu, Wenbo Guo, Xinyu Xing, William Stafford NobleICML 2021 · 14 citations
- Concise Explanations of Neural Networks using Adversarial TrainingPrasad Chalasani, Jiefeng Chen, Amrita Roy Chowdhury, Xi Wu et al.ICML 2020 · 148 citations
- Smoothed Geometry for Robust AttributionZifan Wang, Haofan Wang, Shakul Ramkumar, Piotr Mardziel et al.NeurIPS 2020 · 67 citations
- Robust Explanation Constraints for Neural NetworksMatthew Wicker, Juyeon Heo, Luca Costabello, Adrian WellerICLR 2023 · 3 citations
