Seeing but Not Believing: Probing the Disconnect Between Visual Attention and Answer Correctness in VLMs
Zhining Liu, Ziyi Chen, Hui Liu, Chen Luo, Xianfeng Tang, Suhang Wang, Jingying Zeng, Zhenwei Dai, Zhan Shi, Tianxin Wei, Hanqing Lu, Benoît Dumoulin, Hanghang Tong
Abstract
Vision-Language Models (VLMs) achieve strong results on multimodal tasks such as visual question answering, yet they can still fail even when the correct visual evidence is present. In this work, we systematically investigate whether these failures arise from not perceiving the evidence or from not leveraging it effectively. By examining layer-wise attention dynamics, we find that shallow layers focus primarily on text, while deeper layers sparsely but reliably attend to localized evidence regions. Surprisingly, VLMs often perceive the visual evidence when outputting incorrect answers, a phenomenon we term "seeing but not believing" that widely exists in major VLM families. Building on this, we introduce an inference-time intervention that highlights deep-layer evidence regions through selective attention-based masking. It requires no training and consistently improves accuracy across multiple families, including LLaVA, Qwen, Gemma, and InternVL. These results show that VLMs encode reliable evidence internally but under-utilize it, and that making such signals explicit can bridge the gap between perception and reasoning, advancing the diagnostic understanding and reliability of VLMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 435cbf9e-2c9c-4177-a3c1-fe878dc5cbc1Cited by top-tier papers10
- Vision Language Models are BiasedAn Vo, Khai-Nguyen Nguyen, Mohammad Reza Taesiri, Thi Tuong Vy Dang et al.ICLR 2026 · 68 citations
- Zooming without Zooming: Region-to-Image Distillation for Fine-Grained Multimodal PerceptionLai Wei, Liangbo He, jun lan, Lingzhong Dong et al.ICML 2026 · 27 citations
- On Robustness and Chain-of-Thought Consistency of RL-Finetuned VLMsRosie Zhao, Anshul Shah, Xiaoyu Zhu, Xinke Deng et al.ICML 2026 · 10 citations
- Continual Low-Rank Adapters for LLM-based Generative Recommender SystemsHyunsik Yoo, Ting-Wei Li, SeongKu Kang, Zhining Liu et al.ICLR 2026 · 9 citations
- Do VLMs Perceive or Recall? Probing Visual Perception vs. Memory with Classic Visual IllusionsXiaoxiao Sun, Mingyang Li, Kun Yuan, Min Woo Sun et al.CVPR 2026 · 8 citations
Builds on12
- Is A Picture Worth A Thousand Words? Delving Into Spatial Reasoning for Vision Language ModelsJiayu Wang, Yifei Ming, Zhenmei Shi, Vibhav Vineet et al.NeurIPS 2024 · 166 citations
- Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMsShengbang Tong, Zhuang Liu, Yuexiang Zhai, Yi Ma et al.CVPR 2024 · 111 citations
- Astute RAG: Overcoming Imperfect Retrieval Augmentation and Knowledge Conflicts for Large Language ModelsFei Wang, Xingchen Wan, Ruoxi Sun, Jiefeng Chen et al.ACL 2025 · 50 citations
- Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language ModelsNanxing Hu, Xiaoyue Duan, Jinchao Zhang, Guoliang KangACM MM 2025 · 1 citation
- Why Is Spatial Reasoning Hard for VLMs? An Attention Mechanism Perspective on Focus AreasShiqi Chen, Tongyao Zhu, Ruochen Zhou, Jinghan Zhang et al.ICML 2025 · 1 citation
Related papers
- Deeper Thought, Weaker Aim: Understanding and Mitigating Perceptual Impairment during Reasoning in Multimodal Large Language ModelsRuiying Peng, Xueyu Wu, Jing Lei, Lu Hou et al.CVPR 2026 · 4 citations
- Unveiling the Response of Large Vision-Language Models to Visually Absent TokensSohee Kim, Soohyun Ryu, Joonhyung Park, Eunho YangEMNLP 2025
- Too Late to Recall: Explaining the Two-Hop Problem in Multimodal Knowledge RetrievalConstantin Venhoff, Ashkan Khakzar, Sonia Joseph, Philip H. S. Torr et al.NeurIPS 2025 · 8 citations
- The Geometry of Representational Failures in Vision Language ModelsDaniele Savietto, Declan Campbell, André Panisson, Marco Nurisso et al.ICML 2026 · 5 citations
- Finding the Correct Visual Evidence Without Forgetting: Mitigating Hallucination in LVLMs via Inter-Layer Visual Attention DiscrepancyYutong Xie, Zhenglin Hua, Ran Wang, Wing W. Y. Ng et al.ICML 2026 · 1 citation
