NoiseGrad - Enhancing Explanations by Introducing Stochasticity to Model Weights
Kirill Bykov, Anna Hedström, Shinichi Nakajima, Marina M.-C. Höhne
Abstract
Many efforts have been made for revealing the decision-making process of black-box learning machines such as deep neural networks, resulting in useful local and global explanation methods. For local explanation, stochasticity is known to help: a simple method, called SmoothGrad, has improved the visual quality of gradient-based attribution by adding noise to the input space and averaging the explanations of the noisy inputs. In this paper, we extend this idea and propose NoiseGrad that enhances both local and global explanation methods. Specifically, NoiseGrad introduces stochasticity in the weight parameter space, such that the decision boundary is perturbed. NoiseGrad is expected to enhance the local explanation, similarly to SmoothGrad, due to the dual relationship between the input perturbation and the decision boundary perturbation. We evaluate NoiseGrad and its fusion with SmoothGrad - FusionGrad - qualitatively and quantitatively with several evaluation criteria, and show that our novel approach significantly outperforms the baseline methods. Both NoiseGrad and FusionGrad are method-agnostic and as handy as SmoothGrad using a simple heuristic for the choice of the hyperparameter setting without the need of fine-tuning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4b86b663-3d24-4265-b40d-2be2449e273aCited by top-tier papers7
- SoK: Explainable Machine Learning in Adversarial EnvironmentsMaximilian Noppel, Christian WressneggerS&P 2024 · 28 citations
- Robust Explanation for Free or At the Cost of FaithfulnessZeren Tan, Yang TianICML 2023 · 12 citations
- Provably Better Explanations with Optimized Aggregation of Feature AttributionsThomas Decker, Ananta R. Bhattarai, Jindong Gu, Volker Tresp et al.ICML 2024 · 7 citations
- Beyond Single Path Integrated Gradients for Reliable Input Attribution via Randomized Path SamplingGiyoung Jeon, Haedong Jeong, Jaesik ChoiICCV 2023 · 3 citations
- AdaptGrad: Adaptive Sampling to Reduce NoiseLinjiang Zhou, Chao Ma, Zepeng Wang, Libing Wu et al.NeurIPS 2025 · 3 citations
Builds on3
- How Good is the Bayes Posterior in Deep Neural Networks Really?Florian Wenzel, Kevin Roth, Bastiaan S. Veeling, Jakub Swiatkowski et al.ICML 2020 · 409 citations
- Concise Explanations of Neural Networks using Adversarial TrainingPrasad Chalasani, Jiefeng Chen, Amrita Roy Chowdhury, Xi Wu et al.ICML 2020 · 148 citations
- Estimating Model Uncertainty of Neural Networks in Sparse Information FormJongseok Lee, Matthias Humt, Jianxiang Feng, Rudolph TriebelICML 2020 · 54 citations
Related papers
- Distilled Gradient Aggregation: Purify Features for Input Attribution in the Deep Neural NetworkGiyoung Jeon, Haedong Jeong, Jaesik ChoiNeurIPS 2022 · 11 citations
- Saliency strikes back: How filtering out high frequencies improves white-box explanationsSabine Muzellec, Thomas Fel, Victor Boutin, Léo Andéol et al.ICML 2024 · 4 citations
- Robust Models Are More Interpretable Because Attributions Look NormalZifan Wang, Matt Fredrikson, Anupam DattaICML 2022 · 33 citations
- AttEXplore: Attribution for Explanation with model parameters eXplorationZhiyu Zhu, Huaming Chen, Jiayu Zhang, Xinyi Wang et al.ICLR 2024 · 13 citations
- ABLE: Using Adversarial Pairs to Construct Local Models for Explaining Model PredictionsKrishna Khadka, Sunny Shree, Pujan Budhathoki, Yu Lei et al.KDD 2026
