Rethinking Interpretation: Input-Agnostic Saliency Mapping of Deep Visual Classifiers
Naveed Akhtar, Mohammad Amir Asim Khan Jalwana
Abstract
Saliency methods provide post-hoc model interpretation by attributing input features to the model outputs. Current methods mainly achieve this using a single input sample, thereby failing to answer input-independent inquiries about the model. We also show that input-specific saliency mapping is intrinsically susceptible to misleading feature attribution. Current attempts to use `general' input features for model interpretation assume access to a dataset containing those features, which biases the interpretation. Addressing the gap, we introduce a new perspective of input-agnostic saliency mapping that computationally estimates the high-level features attributed by the model to its outputs. These features are geometrically correlated, and are computed by accumulating model's gradient information with respect to an unrestricted data distribution. To compute these features, we nudge independent data points over the model loss surface towards the local minima associated by a human-understandable concept, e.g., class label for classifiers. With a systematic projection, scaling and refinement process, this information is transformed into an interpretable visualization without compromising its model-fidelity. The visualization serves as a stand-alone qualitative interpretation. With an extensive evaluation, we not only demonstrate successful visualizations for a variety of concepts for large-scale models, but also showcase an interesting utility of this new form of saliency mapping by identifying backdoor signatures in compromised classifiers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4f067057-a5bb-477d-b25a-14344c8d1af9Cited by top-tier papers3
- Interpretable3D: An Ad-Hoc Interpretable Classifier for 3D Point CloudsTuo Feng, Ruijie Quan, Xiaohan Wang, Wenguan Wang et al.AAAI 2024 · 31 citations
- Towards credible visual model interpretation with path attributionNaveed Akhtar, Mohammad A. A. K. JalwanaICML 2023 · 6 citations
- Knowledge-Aware Neuron Interpretation for Scene ClassificationYong Guan, Freddy Lécué, Jiaoyan Chen, Ru Li et al.AAAI 2024 · 3 citations
Builds on4
- Understanding Deep Networks via Extremal Perturbations and Smooth MasksRuth Fong, Mandela Patrick, Andrea VedaldiICCV 2019 · 480 citations
- XRAI: Better Attributions Through RegionsAndrei Kapishnikov, Tolga Bolukbasi, Fernanda B. Viégas, Michael TerryICCV 2019 · 251 citations
- There and Back Again: Revisiting Backpropagation Saliency MethodsSylvestre-Alvise Rebuffi, Ruth Fong, Xu Ji, Andrea VedaldiCVPR 2020
- CAMERAS: Enhanced Resolution and Sanity Preserving Class Activation Mapping for Image SaliencyMohammad A. A. K. Jalwana, Naveed Akhtar, Mohammed Bennamoun, Ajmal MianCVPR 2021
Related papers
- Backdoor Attacks on the DNN Interpretation SystemShihong Fang, Anna ChoromanskaAAAI 2022 · 22 citations
- DANCE: Enhancing saliency maps using decoysYang Young Lu, Wenbo Guo, Xinyu Xing, William Stafford NobleICML 2021 · 14 citations
- What Do You See?: Evaluation of Explainable Artificial Intelligence (XAI) Interpretability through Neural BackdoorsYi-Shan Lin, Wen-Chuan Lee, Z. Berkay CelikKDD 2021 · 62 citations
- Unlearning-based Neural InterpretationsChing Lam Choi, Alexandre Duplessis, Serge J. BelongieICLR 2025
- Xplain: Analyzing Invisible Correlations in Model ExplanationKavita Kumari, Alessandro Pegoraro, Hossein Fereidooni, Ahmad-Reza SadeghiUSENIX Security 2024
