Decomposing Query-Key Feature Interactions Using Contrastive Covariances
Andrew Lee, Yonatan Belinkov, Fernanda Viégas, Martin Wattenberg
摘要
Despite the central role of attention heads in Transformers, we lack tools to understand why a model attends to a particular token. To address this, we study the query-key (QK) space -- the bilinear joint embedding space between queries and keys. We present a contrastive covariance method to decompose the QK space into low-rank, human-interpretable components. It is when features in keys and queries align in these low-rank subspaces that high attention scores are produced. We first study our method both analytically and empirically in a simplified setting. We then apply our method to large language models to identify human-interpretable QK subspaces for categorical semantic features and binding features. Finally, we demonstrate how attention scores can be attributed to our identified features.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- How do Language Models Bind Entities in Context?Jiahai Feng, Jacob SteinhardtICLR 2024 · 被引用 81 次
- Talking Heads: Understanding Inter-Layer Communication in Transformer Language ModelsJack Merullo, Carsten Eickhoff, Ellie PavlickNeurIPS 2024 · 被引用 49 次
- Language Models Use Lookbacks to Track BeliefsNikhil Prakash, Natalie Shapira, Arnab Sen Sharma, Christoph Riedl 等ICLR 2026 · 被引用 42 次
- VL-InterpreT: An Interactive Visualization Tool for Interpreting Vision-Language TransformersEstelle Aflalo, Meng Du, Shao-Yen Tseng, Yongfei Liu 等CVPR 2022 · 被引用 34 次
相关 Paper
- Dissecting Query-Key Interaction in Vision TransformersXu Pan, Aaron Philip, Ziqian Xie, Odelia SchwartzNeurIPS 2024 · 被引用 18 次
- Pinpointing Attention-Causal Communication in Language ModelsGabriel Franco, Mark CrovellaNeurIPS 2025 · 被引用 3 次
- AttentionViz: A Global View of Transformer AttentionCatherine Yeh, Yida Chen, Aoyu Wu, Cynthia Chen 等IEEE VIS 2023 · 被引用 78 次
- Singular Vectors of Attention Heads Align with FeaturesGabriel Franco, Carson Loughridge, Mark CrovellaICML 2026 · 被引用 1 次
- More Identifiable yet Equally Performant Transformers for Text ClassificationRishabh Bhardwaj, Navonil Majumder, Soujanya Poria, Eduard H. HovyACL 2021
