Understanding and Mitigating Hallucination in Large Vision-Language Models via Modular Attribution and Intervention
Tianyun Yang, Ziniu Li, Juan Cao, Chang Xu
摘要
Large Vision-Language Models (LVLMs) exhibit impressive capabilities in complex visual tasks but are prone to hallucination, especially in open-ended generation tasks. This paper explores why LVLMs tend to hallucinate and how to mitigate it. First, we conduct causal mediation analysis through counterfactual edits on specific modules in LVLMs. Our results disclose that Multi-Head Attention (MHA) modules contribute more to the probability of generating hallucination words than multi-layer perceptron modules. We then identify specific heads that are responsible for hallucination, referred to as hallucination heads. Second, we examine the behavior of hallucination heads. We find that they are concentrated in the middle and deeper layers, displaying a strong attention bias toward text tokens. Further, we show that the attention patterns of certain hallucination heads exhibit greater similarity to the base language model and change slowly during the instruction tuning process. Finally, we propose two simple yet effective methods to mitigate hallucination: one is training-free and can be applied directly during decoding, while the other involves fine-tuning. Both methods are targeted for hallucination heads to reduce their reliance on text tokens. Notably, our methods achieve up to 1.7x reduction in hallucination rate for the LLaVA-v1.5-7B model in COCO captioning task, outperforming existing baselines. Overall, our findings suggest that hallucinations in LVLMs are likely to stem from certain modules, and targeted interventions can effectively mitigate these issues. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Intervene-All-Paths: Unified Mitigation of LVLM Hallucinations across Alignment FormatsJiaye Qian, Ge Zheng, Yuchen Zhu, Sibei YangNeurIPS 2025 · 被引用 11 次
- Not All Tokens and Heads Are Equally Important: Dual-Level Attention Intervention for Hallucination MitigationLexiang Tang, Xianwei Zhuang, Bang Yang, Zhiyuan Hu 等AAAI 2026 · 被引用 8 次
- VIB-Probe: Detecting and Mitigating Hallucinations in Vision-Language Models via Variational Information BottleneckFeiran Zhang, Yixin Wu, Zhenghua Wang, Xiaohua Wang 等ACL 2026 · 被引用 7 次
- Tell Model Where to Look: Mitigating Hallucinations in MLLMs by Vision-Guided AttentionJianfei Zhao, Feng Zhang, Xin Sun, Chong Feng 等CVPR 2026 · 被引用 6 次
- KVSmooth: Mitigating Hallucination in Multi-modal Large Language Models through Key-Value SmoothingSiyu Jiang, Feiyang Chen, Xiaojin Zhang, Kun HeCVPR 2026 · 被引用 3 次
它引用的顶会 Paper22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li 等ICLR 2024 · 被引用 3,079 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
相关 Paper
- Devils in Middle Layers of Large Vision-Language Models: Interpreting, Detecting and Mitigating Object Hallucinations via Attention LensZhangqi Jiang, Junkai Chen, Beier Zhu, Tingjin Luo 等CVPR 2025
- Cracking the Code of Hallucination in LVLMs with Vision-aware Head DivergenceJinghan He, Kuan Zhu, Haiyun Guo, Junfeng Fang 等ACL 2025
- AdaIAT: Adaptively Increasing Attention to Generated Text to Alleviate Hallucinations in LVLMLian Zhong, Ziqiang He, Jibin Zheng, Jin Li 等CVPR 2026 · 被引用 2 次
- CausalLens: Sensitivity-Guided Multi-Head Causal Intervention for Hallucination Mitigation in Large Vision-Language ModelsJunyang Ji, Qifan Liu, Wenming Yang, Zhihai HeCVPR 2026
- Imitating the Truth: Attention-aware Truth-Guided Enhancement for Hallucination Mitigation in Large Vision-Language ModelsHairui Ren, Zixuan Wang, Yibo Yang, He Zhao 等ICLR 2026
