DiVE: Decoupling Intra-layer Visual Evidence for Mitigating Hallucinations in Large Vision-Language Models
Xinwei Li, Li Lin, Hui Jiao, Li Yao, Tien-Tsin Wong, Hanqian Wu
Abstract
Recent Large Vision-Language Models (LVLMs) have achieved significant progress yet frequently suffer from visual hallucinations, often stemming from an over-reliance on language priors rather than visual evidence. Existing decoding-based approaches often rely on input perturbations to weaken language priors, but they do not explicitly decouple visual evidence from mixed vision-language representations. To address these limitations, we propose DiVE (Decoupling intra-layer Visual Evidence). DiVE dynamically identifies layers enriched with visual information and performs intra-layer decoupling to extract aggregated visual evidence. By suppressing this evidence to construct a language-priordominated reference distribution, DiVE employs contrastive decoding to calibrate the output logits, thereby mitigating hallucinations. Extensive experiments across diverse LVLM architectures demonstrate that DiVE achieves state-of-the-art performance among decodingbased methods on multiple benchmarks. Crucially, it eliminates the latency of an extra forward pass, offering a lightweight and efficient solution.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7f7cacbd-173a-41d5-afeb-8a4eedd0b92dBuilds on18
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction TuningFuxiao Liu, Kevin Lin, Linjie Li, Jianfeng Wang et al.ICLR 2024 · 476 citations
- DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language ModelsYung-Sung Chuang, Yujia Xie, Hongyin Luo, Yoon Kim et al.ICLR 2024 · 354 citations
- Evaluating Object Hallucination in Large Vision-Language ModelsYifan Li, Yifan Du, Kun Zhou, Jinpeng Wang et al.EMNLP 2023 · 344 citations
Related papers
- DEGAP: Dynamic Entropy-Guided Attention Perturbation for Contrastive Decoding in Large Vision-Language ModelsHyein Seo, Yuna Jeong, Mingyu Kang, Junhyeong Park et al.ICML 2026
- Mitigating Object Hallucinations in Large Vision-Language Models through Visual Contrastive DecodingSicong Leng, Hang Zhang, Guanzheng Chen, Xin Li et al.CVPR 2024
- SECOND: Mitigating Perceptual Hallucination in Vision-Language Models via Selective and Contrastive DecodingWoohyeon Park, Woojin Kim, Jaeik Kim, Jaeyoung DoICML 2025
- Multi-Frequency Contrastive Decoding: Alleviating Hallucinations for Large Vision-Language ModelsBingqian Liu, Fu Zhang, Guoqing Chen, Jingwei ChengEMNLP 2025
- VASparse: Towards Efficient Visual Hallucination Mitigation via Visual-Aware Token SparsificationXianwei Zhuang, Zhihong Zhu, Yuxin Xie, Liming Liang et al.CVPR 2025
