Tell Model Where to Look: Mitigating Hallucinations in MLLMs by Vision-Guided Attention
Jianfei Zhao, Feng Zhang, Xin Sun, Chong Feng, Zhixing Tan
Abstract
Visual attention serves as the primary mechanism through which MLLMs interpret visual information; however, its limited localization capability often leads to hallucinations. We observe that although MLLMs can accurately extract visual semantics from visual tokens, they fail to fully leverage this advantage during subsequent inference. To address this limitation, we propose Vision-Guided Attention (VGA), a training-free method that first constructs precise visual grounding by exploiting the semantic content of visual tokens, and then uses this grounding to guide the model's focus toward relevant visual regions. In image captioning, VGA further refines this guidance dynamically during generation by suppressing regions that have already been described. In VGA, each token undergoes only a single forward pass, introducing a negligible latency overhead. In addition, VGA is fully compatible with efficient attention implementations such as FlashAttention. Extensive experiments across diverse MLLMs and multiple hallucination benchmarks demonstrate that VGA achieves state-of-theart dehallucination performance. Further analysis confirms that explicit visual guidance plays a crucial role in enhancing the visual understanding capabilities of MLLMs. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext eb24a0ae-ec47-4e9b-ad96-2c45b368affaBuilds on23
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Evaluating Object Hallucination in Large Vision-Language ModelsYifan Li, Yifan Du, Kun Zhou, Jinpeng Wang et al.EMNLP 2023 · 344 citations
- Mitigating Hallucination Through Theory-Consistent Symmetric Multimodal Preference OptimizationWenqi Liu, Xuemeng Song, Jiaxi Li, Yinwei Wei et al.NeurIPS 2025 · 18 citations
- Detecting and Mitigating Hallucination in Large Vision Language Models via Fine-Grained AI FeedbackWenyi Xiao, Ziwei Huang, Leilei Gan, Wanggui He et al.AAAI 2025 · 12 citations
- Look Twice Before You Answer: Memory-Space Visual Retracing for Hallucination Mitigation in Multimodal Large Language ModelsXin Zou, Yizhou Wang, Yibo Yan, Yuanhuiyi Lyu et al.ICML 2025 · 1 citation
Related papers
- Imitating the Truth: Attention-aware Truth-Guided Enhancement for Hallucination Mitigation in Large Vision-Language ModelsHairui Ren, Zixuan Wang, Yibo Yang, He Zhao et al.ICLR 2026
- Deeper Thought, Weaker Aim: Understanding and Mitigating Perceptual Impairment during Reasoning in Multimodal Large Language ModelsRuiying Peng, Xueyu Wu, Jing Lei, Lu Hou et al.CVPR 2026 · 4 citations
- Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local AttentionWenbin An, Feng Tian, Sicong Leng, Jiahao Nie et al.CVPR 2025
- Revisit What You See: Revealing Visual Semantics in Vision Tokens to Guide LVLM DecodingBeomsik Cho, Jaehyung KimACL 2026
- Your Large Vision-Language Model Only Needs A Few Attention Heads For Visual GroundingSeil Kang, Jinyeong Kim, Junhyeok Kim, Seong Jae HwangCVPR 2025
