Generate, but Verify: Reducing Hallucination in Vision-Language Models with Retrospective Resampling
Tsung-Han Wu, Heekyung Lee, Jiaxin Ge, Joseph E. Gonzalez, Trevor Darrell, David M. Chan
Abstract
Vision-Language Models (VLMs) excel at visual understanding but often suffer from visual hallucinations, where they generate descriptions of nonexistent objects, actions, or concepts, posing significant risks in safety-critical applications. Existing hallucination mitigation methods typically follow one of two paradigms: generation adjustment, which modifies decoding behavior to align text with visual inputs, and post-hoc verification, where external models assess and correct outputs. While effective, generation adjustment methods often rely on heuristics and lack correction mechanisms, while post-hoc verification is complicated, typically requiring multiple models and tending to reject outputs rather than refine them. In this work, we introduce REVERSE, a unified framework that integrates hallucination-aware training with on-the-fly self-verification. By leveraging a new hallucination-verification dataset containing over 1.3M semi-synthetic samples, along with a novel inference-time retrospective resampling technique, our approach enables VLMs to both detect hallucinations during generation and dynamically revise those hallucinations. Our evaluations show that REVERSE achieves state-of-the-art hallucination reduction, outperforming the best existing methods by up to 12% on CHAIR-MSCOCO and 34% on HaloQuest.
Code Model Checkpoints/Datasets REVERSE (REtrospective VERification and SElf-correction) is a hallucination reduction paradigm for Vision-Language Models (VLMs) that unifies generation adjustment and post-hoc verification methods.
REVERSE allows VLMs to be hallucination-aware by explicitly modeling and monitoring the likelihood that each generated phrase is well-grounded. During training, the model is explicitly trained to classify each groundable phrase as either "confident" or "unconfident" and during inference, the model generates responses while continuously verifying the confidence of each phrase using the likelihood of the "unconfident" predictor. If a phrase is sufficiently ungrounded, the model then performs retrospective adjustment to refine the segment, enabling self-correction on the fly. Key to the first goal of classifying each phrase as "confident" or "unconfident" is training the model to understand if a phrase is well-grounded. While VLMs and LLMs inherently provide implicit confidence scores through token probabilities, these scores are often mis-calibrated and do not consistently correlate with output correctness, making them unreliable for verification [50,15,18]. Furthermore, even when accurate, these probabilities offer no indication of where to backtrack for phrase re-generation and self-correction.
To overcome these limitations, we introduce three tokens to the VLM vocabulary that can be used to explicitly mark key phrases and represent the model's confidence level:
• <SPAN>: Marks the beginning of key or object phrases.
• </CN>: Marks the end of confident, grounded phrases.
• </UN>: Marks the end of unconfident, hallucinated phrases.
These tokens, when placed before/after objects or phrases in the scene can serve as ad-hoc classifiers of the confidence of the model. I.e. if a model generates a </UN> token after a phrase, that phrase can be considered to be ungrounded, while if it generates a </CN>, that phrase is likely grounded in the image. Annotating our data with such tokens, as is shown in Figure 2, will allow us to train the VLM itself to perform post-hoc verification instead of relying on an external model.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 84a05c43-43a9-491a-b9e3-ee91c3d5bb28Cited by top-tier papers10
- Hallucination Begins Where Saliency DropsXiaofeng Zhang, Yuanchao Zhu, Chaochen Gu, Xiaosong Yuan et al.ICLR 2026 · 11 citations
- Vision-Language Introspection: Mitigating Overconfident Hallucinations in MLLMs via Interpretable Bi-Causal SteeringShuliang Liu, Songbo Yang, Dong Fang, Sihang Jia et al.ACL 2026 · 9 citations
- MoD-DPO: Towards Mitigating Cross-modal Hallucinations in Omni LLMs using Modality Decoupled Preference OptimizationAshutosh Chaubey, Jiacheng Pang, Mohammad SoleymaniCVPR 2026 · 7 citations
- InEx: Hallucination Mitigation via Introspection and Cross-Modal Multi-Agent CollaborationZhongyu Yang, Yingfang Yuan, Xuanming Jiang, Baoyi An et al.AAAI 2026 · 5 citations
- FINER: MLLMs Hallucinate under Fine-grained Negative QueriesRui Xiao, Sanghwan Kim, Yongqin Xian, Zeynep Akata et al.CVPR 2026 · 3 citations
Builds on25
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
Related papers
- Self-Correcting Decoding with Generative Feedback for Mitigating Hallucinations in Large Vision-Language ModelsCe Zhang, Zifu Wan, Zhehan Kan, Martin Q. Ma et al.ICLR 2025
- Hallucination-aware Intermediate Representation Edit in Large Vision-Language ModelsWei Suo, Hanzu Zhang, Lijun Zhang, Ji Ma et al.ICLR 2026 · 1 citation
- HalLoc: Token-level Localization of Hallucinations for Vision Language ModelsEunkyu Park, Minyeong Kim, Gunhee KimCVPR 2025
- Antidote: A Unified Framework for Mitigating LVLM Hallucinations in Counterfactual Presupposition and Object PerceptionYuanchen Wu, Lu Zhang, Hang Yao, Junlong Du et al.CVPR 2025
- Anchor-Final Self-Supervision Drives Hallucination-Aware Optimization in Large Vision-Language ModelsJiaxi Liu, Yifeng Yang, Xinbing Wang, Qinying Gu et al.ICML 2026
