Head Pursuit: Probing Attention Specialization in Multimodal Transformers
Lorenzo Basile, Valentino Maiorca, Diego Doimo, Francesco Locatello, Alberto Cazzaniga
摘要
Language and vision-language models have shown impressive performance across a wide range of tasks, but their internal mechanisms remain only partly understood. In this work, we study how individual attention heads in text-generative models specialize in specific semantic or visual attributes. Building on an established interpretability method, we reinterpret the practice of probing intermediate activations with the final decoding layer through the lens of signal processing. This lets us analyze multiple samples in a principled way and rank attention heads based on their relevance to target concepts. Our results show consistent patterns of specialization at the head level across both unimodal and multimodal transformers. Remarkably, we find that editing as few as 1% of the heads, selected using our method, can reliably suppress or enhance targeted concepts in the model output. We validate our approach on language tasks such as question answering and toxicity mitigation, as well as vision-language tasks including image classification and captioning. Our findings highlight an interpretable and controllable structure within attention layers, offering simple tools for understanding and editing large-scale generative models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Towards Understanding Modality Interaction in Multimodal Language Models via Partial Information DecompositionWanlong Fang, Tianle Zhang, Wen Tao, Alvin ChanICML 2026 · 被引用 17 次
- When Seeing Overrides Knowing: Disentangling Knowledge Conflicts in Vision-Language ModelsFrancesco Ortu, Zhijing Jin, Diego Doimo, Alberto CazzanigaACL 2026 · 被引用 7 次
- The Narrow Gate: Localized Image-Text Communication in Native Multimodal ModelsAlessandro Serra, Francesco Ortu, Emanuele Panizon, Lucrezia Valeriani 等NeurIPS 2025 · 被引用 4 次
- Large Vision-Language Models Get Lost in AttentionGongli Xi, Ye Tian, Mengyu Yang, Huahui Yi 等ICML 2026 · 被引用 4 次
- Think Parallax: Solving Multi-Hop Problems via Multi-View Knowledge-Graph-Based Retrieval-Augmented GenerationJinliang Liu, Jiale Bai, Shaoning ZengACL 2026 · 被引用 1 次
它引用的顶会 Paper19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 被引用 3,415 次
- The Linear Representation Hypothesis and the Geometry of Large Language ModelsKiho Park, Yo Joong Choe, Victor VeitchICML 2024 · 被引用 461 次
相关 Paper
- Compressed Sensing for Capability Localization in Large Language ModelsAnna Bair, Yixuan Xu, Mingjie Sun, Zico KolterICML 2026 · 被引用 1 次
- From Weights to Concepts: Data-Free Interpretability of CLIP via Singular Vector DecompositionFrancesco Gentile, Nicola DallAsen, Francesco Tonini, Massimiliano Mancini 等CVPR 2026
- Cross-Attention Head Position Patterns Can Align with Human Visual Concepts in Text-to-Image Generative ModelsJungwon Park, Jungmin Ko, Dongnam Byun, Jangwon Suh 等ICLR 2025
- VIB-Probe: Detecting and Mitigating Hallucinations in Vision-Language Models via Variational Information BottleneckFeiran Zhang, Yixin Wu, Zhenghua Wang, Xiaohua Wang 等ACL 2026 · 被引用 7 次
- V-SEAM: Visual Semantic Editing and Attention Modulating for Causal Interpretability of Vision-Language ModelsQidong Wang, Junjie Hu, Ming JiangEMNLP 2025 · 被引用 2 次
