Lune

ICML2024Top-tier venue

Saliency strikes back: How filtering out high frequencies improves white-box explanations

Sabine Muzellec, Thomas Fel, Victor Boutin, Léo Andéol, Rufin VanRullen, Thomas Serre

2024Year
4Citations
6Top-tier citations

Abstract

Attribution methods correspond to a class of explainability methods (XAI) that aim to assess how individual inputs contribute to a model's decisionmaking process. We have identified a significant limitation in one type of attribution methods, known as "white-box" methods. Although highly efficient, as we will show, these methods rely on a gradient signal that is often contaminated by highfrequency artifacts. To overcome this limitation, we introduce a new approach called "FORGrad." This simple method effectively filters out these high-frequency artifacts using optimal cut-off frequencies tailored to the unique characteristics of each model architecture. Our findings show that FORGrad consistently enhances the performance of existing white-box methods, enabling them to compete effectively with more accurate yet computationally more demanding "black-box" methods. We anticipate that, because of its effectiveness, the proposed method will foster the broader adoption of straightforward and efficient whitebox methods for explainability, providing a better balance between faithfulness and computational efficiency.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

Cited by top-tier papers6

Ask how each one uses it

Builds on15

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines