From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers
Jingtong Su, Julia Kempe, Karen Ullrich
Abstract
Transformers have achieved state-of-the-art performance across diverse language and vision tasks. This success drives the imperative to interpret their internal mechanisms with the dual goals of enhancing performance and improving behavioral control. Attribution methods help advance interpretability by assigning model outputs associated with a target concept to specific model components. Current attribution research primarily studies multi-layer perceptron (MLP) neurons and addresses relatively simple concepts such as factual associations (e.g., Paris is located in France). This focus tends to overlook the impact of the attention mechanism and lacks a unified approach for analyzing more complex concepts. To fill these gaps, we introduce Scalable Attention Module Discovery (SAMD), a concept-agnostic method for mapping arbitrary, complex concepts to specific attention heads of general transformer models. We accomplish this by representing each concept as a vector, calculating its cosine similarity with each attention head, and selecting the TopK-scoring heads to construct the concept-associated attention module. We then propose Scalar Attention Module Intervention (SAMI), a simple strategy to diminish or amplify the effects of a concept by adjusting the attention module using only a single scalar parameter. Empirically, we demonstrate SAMD on concepts of varying complexity, and visualize the locations of their corresponding modules. Our results demonstrate that module locations remain stable before and after LLM post-training, and confirm prior work on the mechanics of LLM multi-lingualism. Through SAMI, we facilitate jailbreaking on HarmBench (+72.7%) by diminishing “safety” and improve performance on the GSM8K benchmark (+1.6%) by amplifying “reasoning”. Lastly, we highlight the domain-agnostic nature of our approach by suppressing the image classification accuracy of vision transformers on ImageNet.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Focusing on Language: Revealing and Exploiting Language Attention Heads in Multilingual Large Language ModelsXin Liu, Qiyang Song, Qihang Zhou, Haichao Du et al.AAAI 2026
- DuetMerging: Synergizing Dynamic and Static Strategies for Mitigating Task Interference in Model MergingYan Li, Guiping Cao, Yaguang Song, Ming Tao et al.CVPR 2026
Builds on31
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 2,317 citations
- Refusal in Language Models Is Mediated by a Single DirectionAndy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka et al.NeurIPS 2024 · 1,166 citations
Related papers
- On the Role of Attention Heads in Large Language Model SafetyZhenhong Zhou, Haiyang Yu, Xinghua Zhang, Rongwu Xu et al.ICLR 2025
- Head Pursuit: Probing Attention Specialization in Multimodal TransformersLorenzo Basile, Valentino Maiorca, Diego Doimo, Francesco Locatello et al.NeurIPS 2025 · 21 citations
- Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 SmallKevin Ro Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris et al.ICLR 2023 · 50 citations
- Causal Attention for Vision-Language TasksXu Yang, Hanwang Zhang, Guojun Qi, Jianfei CaiCVPR 2021
- Attention-based Interpretability with Concept TransformersMattia Rigotti, Christoph Miksovic, Ioana Giurgiu, Thomas Gschwind et al.ICLR 2022 · 77 citations
