Saliency strikes back: How filtering out high frequencies improves white-box explanations
Sabine Muzellec, Thomas Fel, Victor Boutin, Léo Andéol, Rufin VanRullen, Thomas Serre
Abstract
Attribution methods correspond to a class of explainability methods (XAI) that aim to assess how individual inputs contribute to a model's decisionmaking process. We have identified a significant limitation in one type of attribution methods, known as "white-box" methods. Although highly efficient, as we will show, these methods rely on a gradient signal that is often contaminated by highfrequency artifacts. To overcome this limitation, we introduce a new approach called "FORGrad." This simple method effectively filters out these high-frequency artifacts using optimal cut-off frequencies tailored to the unique characteristics of each model architecture. Our findings show that FORGrad consistently enhances the performance of existing white-box methods, enabling them to compete effectively with more accurate yet computationally more demanding "black-box" methods. We anticipate that, because of its effectiveness, the proposed method will foster the broader adoption of straightforward and efficient whitebox methods for explainability, providing a better balance between faithfulness and computational efficiency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Smoothed Differentiation Efficiently Mitigates Shattered Gradients in ExplanationsAdrian Hill, Neal McKee, Johannes Maeß, Stefan Bluecher et al.NeurIPS 2025 · 2 citations
- A Distributional View for Visual Mechanistic Interpretability: KL-Minimal Soft-Constraint PrincipleGuancheng Zhou, Yisi Luo, Zhengfu He, Zhenyu Jin et al.ICML 2026 · 1 citation
- Rethinking Explanation Evaluation Under the Retraining SchemeYi Cai, Thibaud Ardoin, Mayank Gulati, Gerhard WunderAAAI 2026
- GEFA: A General Feature Attribution Framework Using Proxy Gradient EstimationYi Cai, Thibaud Ardoin, Gerhard WunderICML 2025
- Boosting the visual interpretability of CLIP via adversarial fine-tuningShizhan Gong, Haoyu Lei, Qi Dou, Farzan FarniaICLR 2025
Builds on15
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- Reliable Post hoc Explanations: Modeling Uncertainty in ExplainabilityDylan Slack, Anna Hilgard, Sameer Singh, Himabindu LakkarajuNeurIPS 2021 · 240 citations
- Evaluating Explainable AI: Which Algorithmic Explanations Help Users Predict Model Behavior?Peter Hase, Mohit BansalACL 2020 · 216 citations
- Sanity Checks for Saliency MetricsRichard Tomsett, Dan Harborne, Supriyo Chakraborty, Prudhvi Gurram et al.AAAI 2020 · 204 citations
Related papers
- Derivative-Free Diffusion Manifold-Constrained Gradient for Unified XAIWon Jun Kim, Hyungjin Chung, Jaemin Kim, Sangmin Lee et al.CVPR 2025
- NoiseGrad - Enhancing Explanations by Introducing Stochasticity to Model WeightsKirill Bykov, Anna Hedström, Shinichi Nakajima, Marina M.-C. HöhneAAAI 2022 · 43 citations
- On Gradient-like Explanation under a Black-box Setting: When Black-box Explanations Become as Good as White-boxYi Cai, Gerhard WunderICML 2024 · 4 citations
- Distribution-Based Feature Attribution for Explaining the Predictions of Any ClassifierXinpeng Li, Kai Ming TingAAAI 2026
- Generative Perturbation Analysis for Probabilistic Black-Box Anomaly AttributionTsuyoshi Idé, Naoki AbeKDD 2023 · 4 citations
