PAS: Prelim Attention Score for Detecting Object Hallucinations in Large Vision-Language Models
Nhat Hoang, Minh Vu, My T. Thai, Manish Bhattarai
Abstract
Large vision-language models (LVLMs) are powerful, yet they remain unreliable due to object hallucinations. In this work, we show that in many hallucinatory predictions the LVLM effectively ignores the image and instead relies on previously generated output ("prelim") tokens to infer new objects. We quantify this behavior via the mutual information between the image and the predicted object conditioned on the prelim, demonstrating that weak image dependence strongly correlates with hallucination. Building on this finding, we introduce the Prelim Attention Score (PAS), a lightweight, training-free signal computed from attention weights over prelim tokens. PAS requires no additional forward passes and can be computed on the fly during inference. Exploiting this previously overlooked signal, PAS achieves state-of-the-art object-hallucination detection across multiple models and datasets, enabling real-time filtering and intervention.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d456d410-a511-4b5a-a886-f1bdafce60acBuilds on18
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
- Uncertainty Estimation in Autoregressive Structured PredictionAndrey Malinin, Mark J. F. GalesICLR 2021 · 439 citations
- Evaluating Object Hallucination in Large Vision-Language ModelsYifan Li, Yifan Du, Kun Zhou, Jinpeng Wang et al.EMNLP 2023 · 344 citations
- Analyzing and Mitigating Object Hallucination in Large Vision-Language ModelsYiyang Zhou, Chenhang Cui, Jaehong Yoon, Linjun Zhang et al.ICLR 2024 · 316 citations
Related papers
- Hallucinatory Image Tokens: A Training-Free EAZY Approach to Detecting and Mitigating Object Hallucinations in LVLMsLiwei Che, Tony Qingze Liu, Jing Jia, Weiyi Qin et al.ICCV 2025 · 2 citations
- DHCP: Detecting Hallucinations by Cross-modal Attention Pattern in Large Vision-Language ModelsYudong Zhang, Ruobing Xie, Xingwu Sun, Yiqing Huang et al.ACM MM 2025 · 1 citation
- Beyond the Global Scores: Fine-Grained Token Grounding as a Robust Detector of LVLM HallucinationsTuan Dung Nguyen, Minh Khoi Ho, Qi Chen, Yutong Xie et al.CVPR 2026 · 4 citations
- AdaIAT: Adaptively Increasing Attention to Generated Text to Alleviate Hallucinations in LVLMLian Zhong, Ziqiang He, Jibin Zheng, Jin Li et al.CVPR 2026 · 2 citations
- Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local AttentionWenbin An, Feng Tian, Sicong Leng, Jiahao Nie et al.CVPR 2025
