Causal Tracing of Object Representations in Large Vision Language Models: Mechanistic Interpretability and Hallucination Mitigation
Qiming Li, Zekai Ye, Xiaocheng Feng, Weihong Zhong, Weitao Ma, Xiachong Feng
Abstract
Despite the remarkable advancements of Large Vision-Language Models (LVLMs), the mechanistic interpretability remains underexplored. Existing analyses are insufficiently comprehensive and lack examination covering visual and textual tokens, model components, and the full range of layers. This limitation restricts actionable insights to improve the faithfulness of model output and the development of downstream tasks, such as hallucination mitigation. To address this limitation, we introduce Fine-grained Cross-modal Causal Tracing (FCCT) framework, which systematically quantifies the causal effects on visual object perception. FCCT conducts fine-grained analysis covering the full range of visual and textual tokens, three core model components including multi-head self-attention (MHSA), feed-forward networks (FFNs), and hidden states, across all decoder layers. Our analysis is the first to demonstrate that MHSAs of the last token in middle layers play a critical role in aggregating cross-modal information, while FFNs exhibit a three-stage hierarchical progression for the storage and transfer of visual object representations. Building on these insights, we propose Intermediate Representation Injection (IRI), a training-free inference-time technique that reinforces visual object information flow by precisely intervening on cross-modal representations at specific components and layers, thereby enhancing perception and mitigating hallucination. Consistent improvements across five widely used benchmarks and LVLMs demonstrate IRI achieves state-of-the-art performance, while preserving inference speed and other foundational performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 19f3993a-d582-4ff6-8ab3-4dffc2cbe754Cited by top-tier papers5
- Make LVLMs Focus: Context-Aware Attention Modulation for Better Multimodal In-Context LearningYanshu Li, Jianjiang Yang, Ziteng Yang, Bozheng Li et al.AAAI 2026 · 9 citations
- Improving Region Representation Learning from Urban Imagery with Noisy Long-Caption SupervisionYimei Zhang, Guojiang Shen, Kaili Ning, Tongwei Ren et al.AAAI 2026 · 3 citations
- Probing Cross-modal Information Hubs in Audio-Visual LLMsJihoo Jung, Chaeyoung Jung, Ji-Hoon Kim, Joon Son ChungICML 2026 · 1 citation
- Mechanisms of Object Localization in Vision-Language ModelsTimothy Schaumlöffel, Martina G. Vilas, Gemma RoigCVPR 2026 · 1 citation
- Interpreting and Enhancing Emotional Circuits in Large Vision-Language Models via Cross-Modal Information FlowChengsheng Zhang, Chenghao Sun, Zhining Xie, Xinmei TianICML 2026
Builds on15
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- Ferret: Refer and Ground Anything Anywhere at Any GranularityHaoxuan You, Haotian Zhang, Zhe Gan, Xianzhi Du et al.ICLR 2024 · 515 citations
- Evaluating Object Hallucination in Large Vision-Language ModelsYifan Li, Yifan Du, Kun Zhou, Jinpeng Wang et al.EMNLP 2023 · 344 citations
- Alleviating Hallucinations in Large Vision-Language Models through Hallucination-Induced OptimizationXinyu Lyu, Beitao Chen, Lianli Gao, Hengtao Shen et al.NeurIPS 2024 · 59 citations
- Understanding Information Storage and Transfer in Multi-Modal Large Language ModelsSamyadeep Basu, Martin Grayson, Cecily Morrison, Besmira Nushi et al.NeurIPS 2024 · 57 citations
Related papers
- CausalLens: Sensitivity-Guided Multi-Head Causal Intervention for Hallucination Mitigation in Large Vision-Language ModelsJunyang Ji, Qifan Liu, Wenming Yang, Zhihai HeCVPR 2026
- Understanding and Mitigating Hallucination in Large Vision-Language Models via Modular Attribution and InterventionTianyun Yang, Ziniu Li, Juan Cao, Chang XuICLR 2025
- Dismantling Pathological Shortcuts: A Causal Framework for Faithful LVLM DecodingLiu Yu, Can Chen, PING KUANG, Zhikun Feng et al.ICML 2026
- CoFi-Dec: Hallucination-Resistant Decoding via Coarse-to-Fine Generative Feedback in Large Vision-Language ModelsZongsheng Cao, Yangfan He, Anran Liu, Jun Xie et al.ACM MM 2025 · 3 citations
- Devils in Middle Layers of Large Vision-Language Models: Interpreting, Detecting and Mitigating Object Hallucinations via Attention LensZhangqi Jiang, Junkai Chen, Beier Zhu, Tingjin Luo et al.CVPR 2025
