Neural Response Interpretation Through the Lens of Critical Pathways
Ashkan Khakzar, Soroosh Baselizadeh, Saurabh Khanduja, Christian Rupprecht, Seong Tae Kim, Nassir Navab
摘要
Is critical input information encoded in specific sparse pathways within the neural network? In this work, we discuss the problem of identifying these critical pathways and subsequently leverage them for interpreting the network's response to an input. The pruning objective -selecting the smallest group of neurons for which the response remains equivalent to the original network -has been previously proposed for identifying critical pathways. We demonstrate that sparse pathways derived from pruning do not necessarily encode critical input information. To ensure sparse pathways include critical fragments of the encoded input information, we propose pathway selection via neurons' contribution to the response. We proceed to explain how critical pathways can reveal critical input features. We prove that pathways selected via neuron contribution are locally linear (in an ℓ 2 -ball), a property that we use for proposing a feature attribution method: "pathway gradient". We validate our interpretation method using mainstream evaluation experiments. The validation of pathway gradient interpretation method further confirms that selected pathways using neuron contributions correspond to critical input features. The code 1 2 is publicly available.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Fine-Grained Neural Network Explanation by Identifying Input Features with Predictive InformationYang Zhang, Ashkan Khakzar, Yawei Li, Azade Farshad 等NeurIPS 2021 · 被引用 33 次
- FAIRER: Fairness as Decision Rationale AlignmentTianlin Li, Qing Guo, Aishan Liu, Mengnan Du 等ICML 2023 · 被引用 20 次
- Do Explanations Explain? Model Knows BestAshkan Khakzar, Pedram Khorsandi, Rozhin Nobahari, Nassir NavabCVPR 2022 · 被引用 16 次
- IRAD: Implicit Representation-driven Image Resampling against Adversarial AttacksYue Cao, Tianlin Li, Xiaofeng Cao, Ivor W. Tsang 等ICLR 2024 · 被引用 4 次
- Interpreting vision transformers via residual replacement modelJinyeong Kim, Junhyeok Kim, Yumin Shim, Joohyeok Kim 等NeurIPS 2025 · 被引用 4 次
它引用的顶会 Paper5
- The Many Shapley Values for Model ExplanationMukund Sundararajan, Amir NajmiICML 2020 · 被引用 799 次
- Understanding Deep Networks via Extremal Perturbations and Smooth MasksRuth Fong, Mandela Patrick, Andrea VedaldiICCV 2019 · 被引用 480 次
- Restricting the Flow: Information Bottlenecks for AttributionKarl Schulz, Leon Sixt, Federico Tombari, Tim LandgrafICLR 2020 · 被引用 220 次
- Visualizing Deep Networks by Optimizing with Integrated GradientsZhongang Qi, Saeed Khorram, Fuxin LiAAAI 2020 · 被引用 149 次
- When Explanations Lie: Why Many Modified BP Attributions FailLeon Sixt, Maximilian Granz, Tim LandgrafICML 2020 · 被引用 147 次
相关 Paper
- Generalizing Backpropagation for Gradient-Based InterpretabilityKevin Du, Lucas Torroba Hennigen, Niklas Stoehr, Alex Warstadt 等ACL 2023 · 被引用 3 次
- Integrated Directional Gradients: Feature Interaction Attribution for Neural NLP ModelsSandipan Sikdar, Parantapa Bhattacharya, Kieran HeeseACL 2021
- SInGE: Sparsity via Integrated Gradients Estimation of Neuron RelevanceEdouard Yvinec, Arnaud Dapogny, Matthieu Cord, Kevin BaillyNeurIPS 2022 · 被引用 12 次
- SPADE: Sparsity-Guided Debugging for Deep Neural NetworksArshia Soltani Moakhar, Eugenia Iofinova, Elias Frantar, Dan AlistarhICML 2024 · 被引用 2 次
- A Functional Information Perspective on Model InterpretationItai Gat, Nitay Calderon, Roi Reichart, Tamir HazanICML 2022 · 被引用 6 次
