Respect the model: Fine-grained and Robust Explanation with Sharing Ratio Decomposition
Sangyu Han, Yearim Kim, Nojun Kwak
摘要
The truthfulness of existing explanation methods in authentically elucidating the underlying model's decision-making process has been questioned. Existing methods have deviated from faithfully representing the model, thus susceptible to adversarial attacks. To address this, we propose a novel eXplainable AI (XAI) method called SRD (Sharing Ratio Decomposition), which sincerely reflects the model's inference process, resulting in significantly enhanced robustness in our explanations. Different from the conventional emphasis on the neuronal level, we adopt a vector perspective to consider the intricate nonlinear interactions between filters. We also introduce an interesting observation termed Activation-Pattern-Only Prediction (APOP), letting us emphasize the importance of inactive neurons and redefine relevance encapsulating all relevant information including both active and inactive neurons. Our method, SRD, allows for the recursive decomposition of a Pointwise Feature Vector (PFV), providing a high-resolution Effective Receptive Field (ERF) at any layer. * Equal contribution. † Corresponding author. 1 A neuron outputs a scalar, an element in a tensor, by combining the information in its receptive field.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper3
- Concise Explanations of Neural Networks using Adversarial TrainingPrasad Chalasani, Jiefeng Chen, Amrita Roy Chowdhury, Xi Wu 等ICML 2020 · 被引用 148 次
- CRAFT: Concept Recursive Activation FacTorization for ExplainabilityThomas Fel, Agustin Martin Picard, Louis Béthune, Thibaut Boissin 等CVPR 2023
- Shortcomings of Top-Down Randomization-Based Sanity Checks for Evaluations of Deep Neural Network ExplanationsAlexander Binder, Leander Weber, Sebastian Lapuschkin, Grégoire Montavon 等CVPR 2023
相关 Paper
- Logic Rule Guided Attribution with Dynamic AblationJianqiao An, Yuandu Lai, Yahong HanAAAI 2022 · 被引用 4 次
- What Do You See?: Evaluation of Explainable Artificial Intelligence (XAI) Interpretability through Neural BackdoorsYi-Shan Lin, Wen-Chuan Lee, Z. Berkay CelikKDD 2021 · 被引用 62 次
- Effective Optimization of Root Selection Towards Improved Explanation of Deep ClassifiersXin Zhang, Shenghua Zhong, Jianmin JiangACM MM 2024
- Linear Explanations for Individual NeuronsTuomas P. Oikarinen, Tsui-Wei WengICML 2024 · 被引用 18 次
- Minimizing False-Positive Attributions in Explanations of Non-Linear ModelsAnders Gjølbye, Stefan Haufe, Lars Kai HansenNeurIPS 2025 · 被引用 3 次
