B-cos Networks: Alignment is All We Need for Interpretability
Moritz Böhle, Mario Fritz, Bernt Schiele
Abstract
We present a new direction for increasing the interpretability of deep neural networks (DNNs) by promoting weight-input alignment during training. For this, we propose to replace the linear transforms in DNNs by our B-cos transform. As we show, a sequence (network) of such transforms induces a single linear transform that faithfully summarises the full model computations. Moreover, the B-cos transform introduces alignment pressure on the weights during optimisation. As a result, those induced linear transforms become highly interpretable and align with task-relevant features. Importantly, the B-cos transform is designed to be compatible with existing architectures and we show that it can easily be integrated into common models such as VGGs, ResNets, InceptionNets, and DenseNets, whilst maintaining similar performance on ImageNet. The resulting explanations are of high visual quality and perform well under quantitative metrics for interpretability. Code available at github.com/moboehle/B-cos.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers38
- "Help Me Help the AI": Understanding How Explainability Can Support Human-AI InteractionSunnie S. Y. Kim, Elizabeth Anne Watkins, Olga Russakovsky, Ruth Fong et al.CHI 2023 · 178 citations
- Learning Support and Trivial Prototypes for Interpretable Image ClassificationChong Wang, Yuyuan Liu, Yuanhong Chen, Fengbei Liu et al.ICCV 2023 · 50 citations
- FunnyBirds: A Synthetic Vision Dataset for a Part-Based Analysis of Explainable AI MethodsRobin Hesse, Simone Schaub-Meyer, Stefan RothICCV 2023 · 50 citations
- ICICLE: Interpretable Class Incremental Continual LearningDawid Rymarczyk, Joost van de Weijer, Bartosz Zielinski, Bartlomiej TwardowskiICCV 2023 · 35 citations
- PDiscoNet: Semantically consistent part discovery for fine-grained recognitionRobert van der Klis, Stephan Alaniz, Massimiliano Mancini, Cássio Fraga Dantas et al.ICCV 2023 · 26 citations
Builds on4
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Do Input Gradients Highlight Discriminative Features?Harshay Shah, Prateek Jain, Praneeth NetrapalliNeurIPS 2021 · 74 citations
- Rethinking the Role of Gradient-based Attribution Methods for Model InterpretabilitySuraj Srinivas, François FleuretICLR 2021 · 46 citations
- Convolutional Dynamic Alignment Networks for Interpretable ClassificationsMoritz Böhle, Mario Fritz, Bernt SchieleCVPR 2021
Related papers
- B-cosification: Transforming Deep Neural Networks to be Inherently InterpretableShreyash Arya, Sukrut Rao, Moritz Böhle, Bernt SchieleNeurIPS 2024 · 14 citations
- Align Once to Explain: Feature Alignment for Scalable B-cosification of Foundational Vision TransformersRaphael Maser, Siddhartha Gairola, Sukrut Rao, Bernt SchieleCVPR 2026
- Harmonizing the object recognition strategies of deep neural networks with humansThomas Fel, Ivan F. Rodriguez Rodriguez, Drew Linsley, Thomas SerreNeurIPS 2022 · 111 citations
- How does Weight Correlation Affect Generalisation Ability of Deep Neural Networks?Gaojie Jin, Xinping Yi, Liang Zhang, Lijun Zhang et al.NeurIPS 2020 · 5 citations
- SIC: Similarity-Based Interpretable Image Classification with Neural NetworksTom Nuno Wolf, Emre Kavak, Fabian Bongratz, Christian WachingerICCV 2025 · 1 citation
