V-SEAM: Visual Semantic Editing and Attention Modulating for Causal Interpretability of Vision-Language Models
Qidong Wang, Junjie Hu, Ming Jiang
Abstract
Recent advances in causal interpretability have extended from language models to visionlanguage models (VLMs), seeking to reveal their internal mechanisms through input interventions. While textual interventions often target semantics, visual interventions typically rely on coarse pixel-level perturbations, limiting semantic insights on multimodal integration. In this study, we introduce V-SEAM, a novel framework that combines Visual Semantic Editing and Attention Modulating for causal interpretation of VLMs. V-SEAM enables concept-level visual manipulations and identifies attention heads with positive or negative contributions to predictions across three semantic levels: objects, attributes, and relationships. We observe that positive heads are often shared within the same semantic level but vary across levels, while negative heads tend to generalize broadly. Finally, we introduce an automatic method to modulate key head embeddings, demonstrating enhanced performance for both LLAVA and In-structBLIP across three diverse VQA benchmarks. Our data and code are released at: https://github.com/petergit1/V-SEAM .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d6efee08-db04-4c1e-88e1-418581ff3c7eCited by top-tier papers1
Ask how each one uses itBuilds on20
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
Related papers
- Map the Flow: Revealing Hidden Pathways of Information in VideoLLMsMinji Kim, Taekyung Kim, Bohyung HanICLR 2026 · 8 citations
- Head Pursuit: Probing Attention Specialization in Multimodal TransformersLorenzo Basile, Valentino Maiorca, Diego Doimo, Francesco Locatello et al.NeurIPS 2025 · 21 citations
- V-Attack: Targeting Disentangled Value Features for Controllable Adversarial Attacks on LVLMsSen Nie, Jie Zhang, Jianxin Yan, Shiguang Shan et al.CVPR 2026 · 9 citations
- Video-Only ToM: Enhancing Theory of Mind in Multimodal Large Language ModelsSiqi Liu, Xinyang Li, Bochao Zou, Junbao Zhuo et al.CVPR 2026
- VIB-Probe: Detecting and Mitigating Hallucinations in Vision-Language Models via Variational Information BottleneckFeiran Zhang, Yixin Wu, Zhenghua Wang, Xiaohua Wang et al.ACL 2026 · 7 citations
