Inferring Functionality of Attention Heads from their Parameters
Amit Elhelo, Mor Geva
摘要
Attention heads are one of the building blocks of large language models (LLMs). Prior work on investigating their operation mostly focused on analyzing their behavior during inference for specific circuits or tasks. In this work, we seek a comprehensive mapping of the operations they implement in a model. We propose MAPS (Mapping Attention head Param-eterS), an efficient framework that infers the functionality of attention heads from their parameters, without any model training or inference. We showcase the utility of MAPS for answering two types of questions: (a) given a predefined operation, mapping how strongly heads across the model implement it, and (b) given an attention head, inferring its salient functionality. Evaluating MAPS on 20 operations across 6 popular LLMs shows its estimations correlate with the head's outputs during inference and are causally linked to the model's predictions. Moreover, its mappings reveal attention heads of certain operations that were overlooked in previous studies, and valuable insights on function universality and architecture biases in LLMs. Next, we present an automatic pipeline and analysis that leverage MAPS to characterize the salient operations of a given head. Our pipeline produces plausible operation descriptions for most heads, as assessed by human judgment, while revealing diverse operations. We release our code and mappings at https://github.com/amitelhelo/MAPS .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Efficient Hallucination Detection for LLMs Using Uncertainty-Aware Attention HeadsArtem Vazhentsev, Lyudmila Rvanova, Gleb Kuzmin, Ekaterina Fadeeva 等ICML 2026 · 被引用 16 次
- Precise In-Parameter Concept Erasure in Large Language ModelsYoav Gur-Arieh, Clara Suslik, Yihuai Hong, Fazl Barez 等EMNLP 2025 · 被引用 10 次
- Large Vision-Language Models Get Lost in AttentionGongli Xi, Ye Tian, Mengyu Yang, Huahui Yi 等ICML 2026 · 被引用 4 次
- Bilinear representation mitigates reversal curse and enables consistent model editingDong-Kyum Kim, Minsung Kim, Jea Kwon, Nakyeong Yang 等ICLR 2026 · 被引用 1 次
- SVD as a Fast Interpretability Method for TransformersMin Xue, Artur AndrzejakICML 2026
它引用的顶会 Paper17
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley 等ICML 2023 · 被引用 1,822 次
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
- Refusal in Language Models Is Mediated by a Single DirectionAndy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka 等NeurIPS 2024 · 被引用 1,166 次
- Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language ModelsAsma Ghandeharioun, Avi Caciularu, Adam Pearce, Lucas Dixon 等ICML 2024 · 被引用 197 次
- Linearity of Relation Decoding in Transformer Language ModelsEvan Hernandez, Arnab Sen Sharma, Tal Haklay, Kevin Meng 等ICLR 2024 · 被引用 163 次
相关 Paper
- Heads up! Large Language Models Can Perform Tasks Without Your Instruction via Selective Attention Head MaskingSenyu Han, Hongchuan Zeng, Kai Yu, Lu ChenICML 2025
- HeadMap: Locating and Enhancing Knowledge Circuits in LLMsXuehao Wang, Liyuan Wang, Binghuai Lin, Yu ZhangICLR 2025
- Token Alignment Heads: Unveiling Attention's Role in LLM Multilingual TranslationBinbin Liu, Wenhan Han, Feng Chen, Yifan Zhang 等ICLR 2026
- Cognitive Mirrors: Exploring the Diverse Functional Roles of Attention Heads in LLM ReasoningXueqi Ma, Jun Wang, Yanbei Jiang, Sarah M. Erfani 等NeurIPS 2025 · 被引用 5 次
- Interpreting Context Look-ups in Transformers: Investigating Attention-MLP InteractionsClement Neo, Shay B. Cohen, Fazl BarezEMNLP 2024 · 被引用 3 次
