Where do Models go Wrong? Parameter-Space Saliency Maps for Explainability
Roman Levin, Manli Shu, Eitan Borgnia, Furong Huang, Micah Goldblum, Tom Goldstein
摘要
Conventional saliency maps highlight input features to which neural network predictions are highly sensitive. We take a different approach to saliency, in which we identify and analyze the network parameters, rather than inputs, which are responsible for erroneous decisions. We first verify that identified salient parameters are indeed responsible for misclassification by showing that turning these parameters off improves predictions on the associated samples, more than turning off the same number of random or least salient parameters. We further validate the link between salient parameters and network misclassification errors by observing that fine-tuning a small number of the most salient parameters on a single sample results in error correction on other samples which were misclassified for similar reasons -nearest neighbors in the saliency space. After validating our parameter-space saliency maps, we demonstrate that samples which cause similar parameters to malfunction are semantically similar. Further, we introduce an input-space saliency counterpart which reveals how image features cause specific network components to malfunction.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 被引用 1,861 次
- Inverting Gradients - How easy is it to break privacy in federated learning?Jonas Geiping, Hartmut Bauermeister, Hannah Dröge, Michael MoellerNeurIPS 2020 · 被引用 1,822 次
- ASAM: Adaptive Sharpness-Aware Minimization for Scale-Invariant Learning of Deep Neural NetworksJungmin Kwon, Jeongseop Kim, Hyunseo Park, In Kwon ChoiICML 2021 · 被引用 385 次
- Efficient Sharpness-aware Minimization for Improved Training of Neural NetworksJiawei Du, Hanshu Yan, Jiashi Feng, Joey Tianyi Zhou 等ICLR 2022 · 被引用 168 次
- Shared Interest: Measuring Human-AI Alignment to Identify Recurring Patterns in Model BehaviorAngie W. Boggust, Benjamin Hoover, Arvind Satyanarayan, Hendrik StrobeltCHI 2022 · 被引用 51 次
相关 Paper
- Passive attention in artificial neural networks predicts human visual selectivityThomas A. Langlois, H. Charles Zhao, Erin Grant, Ishita Dasgupta 等NeurIPS 2021 · 被引用 19 次
- Valid P-Value for Deep Learning-driven Salient RegionDaiki Miwa, Vo Nguyen Le Duy, Ichiro TakeuchiICLR 2023 · 被引用 3 次
- What Do You See?: Evaluation of Explainable Artificial Intelligence (XAI) Interpretability through Neural BackdoorsYi-Shan Lin, Wen-Chuan Lee, Z. Berkay CelikKDD 2021 · 被引用 62 次
- Building Reliable Explanations of Unreliable Neural Networks: Locally Smoothing Perspective of Model InterpretationDohun Lim, Hyeonseok Lee, Sungchan KimCVPR 2021
- DANCE: Enhancing saliency maps using decoysYang Young Lu, Wenbo Guo, Xinyu Xing, William Stafford NobleICML 2021 · 被引用 14 次
