Dissecting Query-Key Interaction in Vision Transformers
Xu Pan, Aaron Philip, Ziqian Xie, Odelia Schwartz
Abstract
Self-attention in vision transformers is often thought to perform perceptual grouping where tokens attend to other tokens with similar embeddings, which could correspond to semantically similar features of an object. However, attending to dissimilar tokens can be beneficial by providing contextual information. We propose to analyze the query-key interaction by the singular value decomposition of the interaction matrix (i.e. ). We find that in many ViTs, especially those with classification training objectives, early layers attend more to similar tokens, while late layers show increased attention to dissimilar tokens, providing evidence corresponding to perceptual grouping and contextualization, respectively. Many of these interactions between features represented by singular vectors are interpretable and semantic, such as attention between relevant objects, between parts of an object, or between the foreground and background. This offers a novel perspective on interpreting the attention mechanism, which contributes to understanding how transformer models utilize context and salient features when processing images.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2fe97b64-7e2f-4f12-b89d-cd3315f382eeCited by top-tier papers8
- Pinpointing Attention-Causal Communication in Language ModelsGabriel Franco, Mark CrovellaNeurIPS 2025 · 3 citations
- Decomposing Query-Key Feature Interactions Using Contrastive CovariancesAndrew Lee, Yonatan Belinkov, Fernanda Viégas, Martin WattenbergICML 2026 · 1 citation
- Singular Vectors of Attention Heads Align with FeaturesGabriel Franco, Carson Loughridge, Mark CrovellaICML 2026 · 1 citation
- From Weights to Concepts: Data-Free Interpretability of CLIP via Singular Vector DecompositionFrancesco Gentile, Nicola DallAsen, Francesco Tonini, Massimiliano Mancini et al.CVPR 2026
- Diffusion-CAM: Faithful Visual Explanations for dMLLMsHaomin Zuo, Yidi Li, Luoxiao Yang, Xiaofeng ZhangACL 2026
Builds on13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- SimMIM: a Simple Framework for Masked Image ModelingZhenda Xie, Zheng Zhang, Yue Cao, Yutong Lin et al.CVPR 2022 · 1,129 citations
Related papers
- Analyzing Vision Transformers for Image Classification in Class Embedding SpaceMartina G. Vilas, Timothy Schaumlöffel, Gemma RoigNeurIPS 2023 · 43 citations
- Attention Guided CAM: Visual Explanations of Vision Transformer Guided by Self-AttentionSaebom Leem, Hyunseok SeoAAAI 2024 · 40 citations
- A Simple Interpretable Transformer for Fine-Grained Image Classification and AnalysisDipanjyoti Paul, Arpita Chowdhury, Xinqi Xiong, Feng-Ju Chang et al.ICLR 2024 · 27 citations
- Vision Transformers Need More Than RegistersCheng Shi, Yizhou Yu, Sibei YangCVPR 2026 · 17 citations
- Learning Correlation Structures for Vision TransformersManjin Kim, Paul Hongsuck Seo, Cordelia Schmid, Minsu ChoCVPR 2024
