B-cosification: Transforming Deep Neural Networks to be Inherently Interpretable
Shreyash Arya, Sukrut Rao, Moritz Böhle, Bernt Schiele
摘要
B-cos Networks have been shown to be effective for obtaining highly human interpretable explanations of model decisions by architecturally enforcing stronger alignment between inputs and weight. B-cos variants of convolutional networks (CNNs) and vision transformers (ViTs), which primarily replace linear layers with B-cos transformations, perform competitively to their respective standard variants while also yielding explanations that are faithful by design. However, it has so far been necessary to train these models from scratch, which is increasingly infeasible in the era of large, pre-trained foundation models. In this work, inspired by the architectural similarities in standard DNNs and B-cos networks, we propose 'B-cosification', a novel approach to transform existing pre-trained models to become inherently interpretable. We perform a thorough study of design choices to perform this conversion, both for convolutional neural networks and vision transformers. We find that B-cosification can yield models that are on par with B-cos models trained from scratch in terms of interpretability, while often outperforming them in terms of classification performance at a fraction of the training cost. Subsequently, we apply B-cosification to a pretrained CLIP model, and show that, even with limited data and compute cost, we obtain a B-cosified version that is highly interpretable and competitive on zero shot performance across a variety of datasets. We release our code and pre-trained model weights at https://github.com/shrebox/B-cosification.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Explaining Similarity in Vision-Language Encoders with Weighted Banzhaf InteractionsHubert Baniecki, Maximilian Muschalik, Fabian Fumagalli, Barbara Hammer 等NeurIPS 2025 · 被引用 6 次
- DAVE: Distribution-aware Attribution via ViT Gradient DecompositionAdam Wróbel, Siddhartha Gairola, Jacek Tabor, Bernt Schiele 等ICML 2026 · 被引用 2 次
- FaCT: Faithful Concept Traces for Explaining Neural Network DecisionsAmin Parchami-Araghi, Sukrut Rao, Jonas Fischer, Bernt SchieleNeurIPS 2025 · 被引用 1 次
- From Pixels to Perception: Interpretable Predictions via Instance-wise Grouped Feature SelectionMoritz Vandenhirtz, Julia E. VogtICML 2025
- Align Once to Explain: Feature Alignment for Scalable B-cosification of Foundational Vision TransformersRaphael Maser, Siddhartha Gairola, Sukrut Rao, Bernt SchieleCVPR 2026
它引用的顶会 Paper25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
- Concept Bottleneck ModelsPang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann 等ICML 2020 · 被引用 1,233 次
- Early Convolutions Help Transformers See BetterTete Xiao, Mannat Singh, Eric Mintun, Trevor Darrell 等NeurIPS 2021 · 被引用 974 次
相关 Paper
- B-cos Networks: Alignment is All We Need for InterpretabilityMoritz Böhle, Mario Fritz, Bernt SchieleCVPR 2022 · 被引用 62 次
- INViTE: INterpret and Control Vision-Language Models with Text ExplanationsHaozhe Chen, Junfeng Yang, Carl Vondrick, Chengzhi MaoICLR 2024 · 被引用 9 次
- Interpreting and Analysing CLIP's Zero-Shot Image Classification via Mutual KnowledgeFawaz Sammani, Nikos DeligiannisNeurIPS 2024 · 被引用 15 次
- Text-To-Concept (and Back) via Cross-Model AlignmentMazda Moayeri, Keivan Rezaei, Maziar Sanjabi, Soheil FeiziICML 2023 · 被引用 62 次
- Synergy and Diversity in CLIP: Enhancing Performance Through Adaptive Backbone EnsemblingCristian Rodriguez Opazo, Ehsan Abbasnejad, Damien Teney, Hamed Damirchi 等ICLR 2025
