Cracking the Code of Hallucination in LVLMs with Vision-aware Head Divergence
Jinghan He, Kuan Zhu, Haiyun Guo, Junfeng Fang, Zhenglin Hua, Yuheng Jia, Ming Tang, Tat-Seng Chua, Jinqiao Wang
Abstract
Large vision-language models (LVLMs) have made substantial progress in integrating large language models (LLMs) with visual inputs, enabling advanced multimodal reasoning. Despite their success, a persistent challenge is hallucination-where generated text fails to accurately reflect visual content-undermining both accuracy and reliability. Existing methods focus on alignment training or decoding refinements but primarily address symptoms at the generation stage without probing the underlying causes. In this work, we investigate the internal mechanisms driving hallucination in LVLMs, with an emphasis on the multi-head attention module. Specifically, we introduce Vision-aware Head Divergence (VHD), a metric that quantifies the sensitivity of attention head outputs to visual context. Based on this, our findings reveal the presence of vision-aware attention heads that are more attuned to visual information; however, the model's overreliance on its prior language patterns is closely related to hallucinations. Building on these insights, we propose Vision-aware Head Reinforcement (VHR), a training-free approach to mitigate hallucination by enhancing the role of visionaware attention heads. Extensive experiments demonstrate that our method achieves superior performance compared to state-of-the-art approaches in mitigating hallucinations, while maintaining high efficiency with negligible additional time overhead. The code is available at https://github.com/jinghan1he/VHR .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a390ad8b-0756-449c-a262-93b818ab78f2Cited by top-tier papers16
- Vision-Language Introspection: Mitigating Overconfident Hallucinations in MLLMs via Interpretable Bi-Causal SteeringShuliang Liu, Songbo Yang, Dong Fang, Sihang Jia et al.ACL 2026 · 9 citations
- VIB-Probe: Detecting and Mitigating Hallucinations in Vision-Language Models via Variational Information BottleneckFeiran Zhang, Yixin Wu, Zhenghua Wang, Xiaohua Wang et al.ACL 2026 · 7 citations
- When Seeing Overrides Knowing: Disentangling Knowledge Conflicts in Vision-Language ModelsFrancesco Ortu, Zhijing Jin, Diego Doimo, Alberto CazzanigaACL 2026 · 7 citations
- Tell Model Where to Look: Mitigating Hallucinations in MLLMs by Vision-Guided AttentionJianfei Zhao, Feng Zhang, Xin Sun, Chong Feng et al.CVPR 2026 · 6 citations
- Finding the Correct Visual Evidence Without Forgetting: Mitigating Hallucination in LVLMs via Inter-Layer Visual Attention DiscrepancyYutong Xie, Zhenglin Hua, Ran Wang, Wing W. Y. Ng et al.ICML 2026 · 1 citation
Builds on16
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister et al.NeurIPS 2023 · 1,549 citations
- DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language ModelsYung-Sung Chuang, Yujia Xie, Hongyin Luo, Yoon Kim et al.ICLR 2024 · 354 citations
- Evaluating Object Hallucination in Large Vision-Language ModelsYifan Li, Yifan Du, Kun Zhou, Jinpeng Wang et al.EMNLP 2023 · 344 citations
- Analyzing and Mitigating Object Hallucination in Large Vision-Language ModelsYiyang Zhou, Chenhang Cui, Jaehong Yoon, Linjun Zhang et al.ICLR 2024 · 316 citations
Related papers
- Understanding and Mitigating Hallucination in Large Vision-Language Models via Modular Attribution and InterventionTianyun Yang, Ziniu Li, Juan Cao, Chang XuICLR 2025
- Not All Tokens and Heads Are Equally Important: Dual-Level Attention Intervention for Hallucination MitigationLexiang Tang, Xianwei Zhuang, Bang Yang, Zhiyuan Hu et al.AAAI 2026 · 8 citations
- DHCP: Detecting Hallucinations by Cross-modal Attention Pattern in Large Vision-Language ModelsYudong Zhang, Ruobing Xie, Xingwu Sun, Yiqing Huang et al.ACM MM 2025 · 1 citation
- CLAIM: Mitigating Multilingual Object Hallucination in Large Vision-Language Models with Cross-Lingual Attention InterventionZekai Ye, Qiming Li, Xiaocheng Feng, Libo Qin et al.ACL 2025 · 14 citations
- Collaboration Wins More: Dual-Modal Collaborative Attention Reinforcement for Mitigating Large Vision Language Models HallucinationJiye Xie, Yifei Gao, Liangliang You, Xiang Xu et al.ACM MM 2025 · 2 citations
