Saliency strikes back: How filtering out high frequencies improves white-box explanations
Sabine Muzellec, Thomas Fel, Victor Boutin, Léo Andéol, Rufin VanRullen, Thomas Serre
摘要
Attribution methods correspond to a class of explainability methods (XAI) that aim to assess how individual inputs contribute to a model's decisionmaking process. We have identified a significant limitation in one type of attribution methods, known as "white-box" methods. Although highly efficient, as we will show, these methods rely on a gradient signal that is often contaminated by highfrequency artifacts. To overcome this limitation, we introduce a new approach called "FORGrad." This simple method effectively filters out these high-frequency artifacts using optimal cut-off frequencies tailored to the unique characteristics of each model architecture. Our findings show that FORGrad consistently enhances the performance of existing white-box methods, enabling them to compete effectively with more accurate yet computationally more demanding "black-box" methods. We anticipate that, because of its effectiveness, the proposed method will foster the broader adoption of straightforward and efficient whitebox methods for explainability, providing a better balance between faithfulness and computational efficiency.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Smoothed Differentiation Efficiently Mitigates Shattered Gradients in ExplanationsAdrian Hill, Neal McKee, Johannes Maeß, Stefan Bluecher 等NeurIPS 2025 · 被引用 2 次
- A Distributional View for Visual Mechanistic Interpretability: KL-Minimal Soft-Constraint PrincipleGuancheng Zhou, Yisi Luo, Zhengfu He, Zhenyu Jin 等ICML 2026 · 被引用 1 次
- Rethinking Explanation Evaluation Under the Retraining SchemeYi Cai, Thibaud Ardoin, Mayank Gulati, Gerhard WunderAAAI 2026
- GEFA: A General Feature Attribution Framework Using Proxy Gradient EstimationYi Cai, Thibaud Ardoin, Gerhard WunderICML 2025
- Boosting the visual interpretability of CLIP via adversarial fine-tuningShizhan Gong, Haoyu Lei, Qi Dou, Farzan FarniaICLR 2025
它引用的顶会 Paper15
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
- Reliable Post hoc Explanations: Modeling Uncertainty in ExplainabilityDylan Slack, Anna Hilgard, Sameer Singh, Himabindu LakkarajuNeurIPS 2021 · 被引用 240 次
- Evaluating Explainable AI: Which Algorithmic Explanations Help Users Predict Model Behavior?Peter Hase, Mohit BansalACL 2020 · 被引用 216 次
- Sanity Checks for Saliency MetricsRichard Tomsett, Dan Harborne, Supriyo Chakraborty, Prudhvi Gurram 等AAAI 2020 · 被引用 204 次
相关 Paper
- Derivative-Free Diffusion Manifold-Constrained Gradient for Unified XAIWon Jun Kim, Hyungjin Chung, Jaemin Kim, Sangmin Lee 等CVPR 2025
- NoiseGrad - Enhancing Explanations by Introducing Stochasticity to Model WeightsKirill Bykov, Anna Hedström, Shinichi Nakajima, Marina M.-C. HöhneAAAI 2022 · 被引用 43 次
- On Gradient-like Explanation under a Black-box Setting: When Black-box Explanations Become as Good as White-boxYi Cai, Gerhard WunderICML 2024 · 被引用 4 次
- Distribution-Based Feature Attribution for Explaining the Predictions of Any ClassifierXinpeng Li, Kai Ming TingAAAI 2026
- Generative Perturbation Analysis for Probabilistic Black-Box Anomaly AttributionTsuyoshi Idé, Naoki AbeKDD 2023 · 被引用 4 次
