On the Complexity-Faithfulness Trade-Off of Gradient-Based Explanations
Amir Mehrpanah, Matteo Gamba, Kevin Smith, Hossein Azizpour
Abstract
ReLU networks, while prevalent for visual data, have sharp transitions, sometimes relying on individual pixels for predictions, making vanilla gradient-based explanations noisy and difficult to interpret. Existing methods, such as GradCAM, smooth these explanations by producing surrogate models at the cost of faithfulness. We introduce a unifying spectral framework to systematically analyze and quantify smoothness, faithfulness, and their trade-off in explanations. Using this framework, we quantify and regularize the contribution of ReLU networks to high-frequency information, providing a principled approach to identifying this trade-off. Our analysis characterizes how surrogate-based smoothing distorts explanations, leading to an “explanation gap” that we formally define and measure for different posthoc methods. Finally, we validate our theoretical findings across different design choices, datasets, and ablations. 11https://github.com/Amir-Mehrpanah/On-the-Complexity-Faithfulness-Trade-off-of-Gradient-Based-Explanations-ICCV25/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 397ab26a-293b-4859-93dc-2b1bcc2ba3e2Cited by top-tier papers1
Ask how each one uses itBuilds on11
- Generalization in diffusion models arises from geometry-adaptive harmonic representationsZahra Kadkhodaie, Florentin Guth, Eero P. Simoncelli, Stéphane MallatICLR 2024 · 168 citations
- Feature Importance Ranking for Deep LearningMaksymilian Wojtas, Ke ChenNeurIPS 2020 · 159 citations
- Which Explanation Should I Choose? A Function Approximation Perspective to Characterizing Post Hoc ExplanationsTessa Han, Suraj Srinivas, Himabindu LakkarajuNeurIPS 2022 · 126 citations
- On the Similarity between the Laplace and Neural Tangent KernelsAmnon Geifman, Abhay Kumar Yadav, Yoni Kasten, Meirav Galun et al.NeurIPS 2020 · 118 citations
- B-cos Networks: Alignment is All We Need for InterpretabilityMoritz Böhle, Mario Fritz, Bernt SchieleCVPR 2022 · 62 citations
Related papers
- Empowering CAM-Based Methods with Capability to Generate Fine-Grained and High-Faithfulness ExplanationsChangqing Qiu, Fusheng Jin, Yining ZhangAAAI 2024 · 11 citations
- Initialization Noise in Image Gradients and Saliency MapsAnn-Christin Woerl, Jan Disselhoff, Michael WandCVPR 2023
- Building Reliable Explanations of Unreliable Neural Networks: Locally Smoothing Perspective of Model InterpretationDohun Lim, Hyeonseok Lee, Sungchan KimCVPR 2021
- Rethinking Attention-Model Explainability through Faithfulness Violation TestYibing Liu, Haoliang Li, Yangyang Guo, Chenqi Kong et al.ICML 2022 · 60 citations
- Measuring the (Un)Faithfulness of Concept-Based ExplanationsShubham Kumar, Narendra AhujaCVPR 2026 · 1 citation
