Neural Response Interpretation Through the Lens of Critical Pathways
Ashkan Khakzar, Soroosh Baselizadeh, Saurabh Khanduja, Christian Rupprecht, Seong Tae Kim, Nassir Navab
Abstract
Is critical input information encoded in specific sparse pathways within the neural network? In this work, we discuss the problem of identifying these critical pathways and subsequently leverage them for interpreting the network's response to an input. The pruning objective -selecting the smallest group of neurons for which the response remains equivalent to the original network -has been previously proposed for identifying critical pathways. We demonstrate that sparse pathways derived from pruning do not necessarily encode critical input information. To ensure sparse pathways include critical fragments of the encoded input information, we propose pathway selection via neurons' contribution to the response. We proceed to explain how critical pathways can reveal critical input features. We prove that pathways selected via neuron contribution are locally linear (in an ℓ 2 -ball), a property that we use for proposing a feature attribution method: "pathway gradient". We validate our interpretation method using mainstream evaluation experiments. The validation of pathway gradient interpretation method further confirms that selected pathways using neuron contributions correspond to critical input features. The code 1 2 is publicly available.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e0e5cd83-c58c-481c-ba87-91ab1d6ea41cCited by top-tier papers12
- Fine-Grained Neural Network Explanation by Identifying Input Features with Predictive InformationYang Zhang, Ashkan Khakzar, Yawei Li, Azade Farshad et al.NeurIPS 2021 · 33 citations
- FAIRER: Fairness as Decision Rationale AlignmentTianlin Li, Qing Guo, Aishan Liu, Mengnan Du et al.ICML 2023 · 20 citations
- Do Explanations Explain? Model Knows BestAshkan Khakzar, Pedram Khorsandi, Rozhin Nobahari, Nassir NavabCVPR 2022 · 16 citations
- IRAD: Implicit Representation-driven Image Resampling against Adversarial AttacksYue Cao, Tianlin Li, Xiaofeng Cao, Ivor W. Tsang et al.ICLR 2024 · 4 citations
- Interpreting vision transformers via residual replacement modelJinyeong Kim, Junhyeok Kim, Yumin Shim, Joohyeok Kim et al.NeurIPS 2025 · 4 citations
Builds on5
- The Many Shapley Values for Model ExplanationMukund Sundararajan, Amir NajmiICML 2020 · 799 citations
- Understanding Deep Networks via Extremal Perturbations and Smooth MasksRuth Fong, Mandela Patrick, Andrea VedaldiICCV 2019 · 480 citations
- Restricting the Flow: Information Bottlenecks for AttributionKarl Schulz, Leon Sixt, Federico Tombari, Tim LandgrafICLR 2020 · 220 citations
- Visualizing Deep Networks by Optimizing with Integrated GradientsZhongang Qi, Saeed Khorram, Fuxin LiAAAI 2020 · 149 citations
- When Explanations Lie: Why Many Modified BP Attributions FailLeon Sixt, Maximilian Granz, Tim LandgrafICML 2020 · 147 citations
Related papers
- Generalizing Backpropagation for Gradient-Based InterpretabilityKevin Du, Lucas Torroba Hennigen, Niklas Stoehr, Alex Warstadt et al.ACL 2023 · 3 citations
- Integrated Directional Gradients: Feature Interaction Attribution for Neural NLP ModelsSandipan Sikdar, Parantapa Bhattacharya, Kieran HeeseACL 2021
- SInGE: Sparsity via Integrated Gradients Estimation of Neuron RelevanceEdouard Yvinec, Arnaud Dapogny, Matthieu Cord, Kevin BaillyNeurIPS 2022 · 12 citations
- SPADE: Sparsity-Guided Debugging for Deep Neural NetworksArshia Soltani Moakhar, Eugenia Iofinova, Elias Frantar, Dan AlistarhICML 2024 · 2 citations
- A Functional Information Perspective on Model InterpretationItai Gat, Nitay Calderon, Roi Reichart, Tamir HazanICML 2022 · 6 citations
