Attentive Multi-Layer Fusion for Vision Transformers
Laure Ciernik, Marco Morik, Lukas Thede, Luca Eyring, Shinichi Nakajima, Zeynep Akata, Lukas Muttenthaler
Abstract
With the rise of large-scale foundation models, efficiently adapting them to downstream tasks remains a central challenge. Linear probing, which freezes the backbone and trains a lightweight head, is computationally efficient but often restricted to last-layer representations. We show that task-relevant information is distributed across the network hierarchy rather than encoded solely in the last layers. To leverage this distribution of information, we apply an attentive probing mechanism that dynamically fuses representations from all layers of a Vision Transformer. This attentive layer fusion (ALF) learns to identify the most relevant layers for a target task and combines low-level structural cues with high-level semantic abstractions. Across 20 diverse datasets and multiple pretrained foundation models, ALF achieves consistent, substantial gains over standard linear probes. Attention heatmaps further reveal that tasks different from the pre-training domain benefit most from intermediate representations. Overall, our findings underscore the value of intermediate layers and demonstrate a principled, task-aware approach for unlocking their potential for probing-based adaptation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0e9488aa-e50a-4606-be2d-fcb035cf4f73Builds on19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image ClassificationChun-Fu (Richard) Chen, Quanfu Fan, Rameswar PandaICCV 2021 · 2,072 citations
- Multiscale Vision TransformersHaoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li et al.ICCV 2021 · 1,611 citations
Related papers
- Fantastic Features and Where to Find Them: A Probing Method to combine Features from Multiple Foundation ModelsBenjamin Ramtoula, Pierre-Yves Lajoie, Paul Newman, Daniele De MartiniNeurIPS 2025 · 2 citations
- Multimodal Language Models See Better When They Look ShallowerHaoran Chen, Junyan Lin, Xinghao Chen, Yue Fan et al.EMNLP 2025
- Visual Query Tuning: Towards Effective Usage of Intermediate Representations for Parameter and Memory Efficient Transfer LearningCheng-Hao Tu, Zheda Mai, Wei-Lun ChaoCVPR 2023
- Head2Toe: Utilizing Intermediate Representations for Better Transfer LearningUtku Evci, Vincent Dumoulin, Hugo Larochelle, Michael C. MozerICML 2022 · 103 citations
- DeAR: Fine-Grained VLM Adaptation by Decomposing Attention Head RolesYiming Ma, Hongkun Yang, Lionel Z. Wang, Bin Chen et al.CVPR 2026 · 2 citations
