B-cos Networks: Alignment is All We Need for Interpretability
Moritz Böhle, Mario Fritz, Bernt Schiele
摘要
We present a new direction for increasing the interpretability of deep neural networks (DNNs) by promoting weight-input alignment during training. For this, we propose to replace the linear transforms in DNNs by our B-cos transform. As we show, a sequence (network) of such transforms induces a single linear transform that faithfully summarises the full model computations. Moreover, the B-cos transform introduces alignment pressure on the weights during optimisation. As a result, those induced linear transforms become highly interpretable and align with task-relevant features. Importantly, the B-cos transform is designed to be compatible with existing architectures and we show that it can easily be integrated into common models such as VGGs, ResNets, InceptionNets, and DenseNets, whilst maintaining similar performance on ImageNet. The resulting explanations are of high visual quality and perform well under quantitative metrics for interpretability. Code available at github.com/moboehle/B-cos.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper38
- "Help Me Help the AI": Understanding How Explainability Can Support Human-AI InteractionSunnie S. Y. Kim, Elizabeth Anne Watkins, Olga Russakovsky, Ruth Fong 等CHI 2023 · 被引用 178 次
- Learning Support and Trivial Prototypes for Interpretable Image ClassificationChong Wang, Yuyuan Liu, Yuanhong Chen, Fengbei Liu 等ICCV 2023 · 被引用 50 次
- FunnyBirds: A Synthetic Vision Dataset for a Part-Based Analysis of Explainable AI MethodsRobin Hesse, Simone Schaub-Meyer, Stefan RothICCV 2023 · 被引用 50 次
- ICICLE: Interpretable Class Incremental Continual LearningDawid Rymarczyk, Joost van de Weijer, Bartosz Zielinski, Bartlomiej TwardowskiICCV 2023 · 被引用 35 次
- PDiscoNet: Semantically consistent part discovery for fine-grained recognitionRobert van der Klis, Stephan Alaniz, Massimiliano Mancini, Cássio Fraga Dantas 等ICCV 2023 · 被引用 26 次
它引用的顶会 Paper4
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Do Input Gradients Highlight Discriminative Features?Harshay Shah, Prateek Jain, Praneeth NetrapalliNeurIPS 2021 · 被引用 74 次
- Rethinking the Role of Gradient-based Attribution Methods for Model InterpretabilitySuraj Srinivas, François FleuretICLR 2021 · 被引用 46 次
- Convolutional Dynamic Alignment Networks for Interpretable ClassificationsMoritz Böhle, Mario Fritz, Bernt SchieleCVPR 2021
相关 Paper
- B-cosification: Transforming Deep Neural Networks to be Inherently InterpretableShreyash Arya, Sukrut Rao, Moritz Böhle, Bernt SchieleNeurIPS 2024 · 被引用 14 次
- Align Once to Explain: Feature Alignment for Scalable B-cosification of Foundational Vision TransformersRaphael Maser, Siddhartha Gairola, Sukrut Rao, Bernt SchieleCVPR 2026
- Harmonizing the object recognition strategies of deep neural networks with humansThomas Fel, Ivan F. Rodriguez Rodriguez, Drew Linsley, Thomas SerreNeurIPS 2022 · 被引用 111 次
- How does Weight Correlation Affect Generalisation Ability of Deep Neural Networks?Gaojie Jin, Xinping Yi, Liang Zhang, Lijun Zhang 等NeurIPS 2020 · 被引用 5 次
- SIC: Similarity-Based Interpretable Image Classification with Neural NetworksTom Nuno Wolf, Emre Kavak, Fabian Bongratz, Christian WachingerICCV 2025 · 被引用 1 次
