Decomposing Query-Key Feature Interactions Using Contrastive Covariances
Andrew Lee, Yonatan Belinkov, Fernanda Viégas, Martin Wattenberg
Abstract
Despite the central role of attention heads in Transformers, we lack tools to understand why a model attends to a particular token. To address this, we study the query-key (QK) space -- the bilinear joint embedding space between queries and keys. We present a contrastive covariance method to decompose the QK space into low-rank, human-interpretable components. It is when features in keys and queries align in these low-rank subspaces that high attention scores are produced. We first study our method both analytically and empirically in a simplified setting. We then apply our method to large language models to identify human-interpretable QK subspaces for categorical semantic features and binding features. Finally, we demonstrate how attention scores can be attributed to our identified features.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0135f2af-2c62-45bc-ae6e-7b94fc44a8adBuilds on14
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- How do Language Models Bind Entities in Context?Jiahai Feng, Jacob SteinhardtICLR 2024 · 81 citations
- Talking Heads: Understanding Inter-Layer Communication in Transformer Language ModelsJack Merullo, Carsten Eickhoff, Ellie PavlickNeurIPS 2024 · 49 citations
- Language Models Use Lookbacks to Track BeliefsNikhil Prakash, Natalie Shapira, Arnab Sen Sharma, Christoph Riedl et al.ICLR 2026 · 42 citations
- VL-InterpreT: An Interactive Visualization Tool for Interpreting Vision-Language TransformersEstelle Aflalo, Meng Du, Shao-Yen Tseng, Yongfei Liu et al.CVPR 2022 · 34 citations
Related papers
- Dissecting Query-Key Interaction in Vision TransformersXu Pan, Aaron Philip, Ziqian Xie, Odelia SchwartzNeurIPS 2024 · 18 citations
- Pinpointing Attention-Causal Communication in Language ModelsGabriel Franco, Mark CrovellaNeurIPS 2025 · 3 citations
- AttentionViz: A Global View of Transformer AttentionCatherine Yeh, Yida Chen, Aoyu Wu, Cynthia Chen et al.IEEE VIS 2023 · 78 citations
- Singular Vectors of Attention Heads Align with FeaturesGabriel Franco, Carson Loughridge, Mark CrovellaICML 2026 · 1 citation
- More Identifiable yet Equally Performant Transformers for Text ClassificationRishabh Bhardwaj, Navonil Majumder, Soujanya Poria, Eduard H. HovyACL 2021
