Lune

CVPR2026顶会

Align Once to Explain: Feature Alignment for Scalable B-cosification of Foundational Vision Transformers

Raphael Maser, Siddhartha Gairola, Sukrut Rao, Bernt Schiele

出版方
2026年份

摘要

Foundational vision models have become the de facto standard for many vision tasks due to their strong performance. However, they are notoriously opaque and remain hard to interpret. We present ALOE (ALign Once to Explain), a one-time, label-free feature alignment based approach that efficiently converts foundational vision models into inherently interpretable B-cos variants. Once aligned, the B-cos backbone is used as a drop-in replacement across several downstream tasks—amortizing the cost of interpretability. ALOE is robust across pre-training paradigms (supervised, self-supervised, vision–language) and is 100–1000× more data-efficient than training from scratch. On classification, it outperforms fully-supervised B-cos models (e.g., +6.6 p.p. top-1 on ImageNet for ViT-B/16), retains strong linear probing, k-NN, and zero-shot transfer performance competitive with foundational backbones (DINOv3, SigLIP2) across diverse downstream datasets, while yielding well-localized and highly human interpretable explanations by design. Code and models will be released.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext 584862ea-e0f4-487b-9564-afc87713c2a2

它引用的顶会 Paper22

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖