Enhancing Vision-Language Model Reliability with Uncertainty-Guided Dropout Decoding
Yixiong Fang, Ziran Yang, Zhaorun Chen, Zhuokai Zhao, Jiawei Zhou
Abstract
Large vision-language models (LVLMs) excel at multimodal tasks but are prone to misinterpreting visual inputs, often resulting in hallucinations and unreliable outputs. We present DROPOUT DECODING, a novel inference-time approach that quantifies the uncertainty of visual tokens and selectively masks uncertain tokens to improve decoding. Our method measures the uncertainty of each visual token by projecting it onto the text space and decomposing it into aleatoric and epistemic components. Specifically, we focus on epistemic uncertainty, which captures perception-related errors more effectively. Inspired by dropout regularization, we introduce uncertainty-guided token dropout, which applies the dropout principle to input visual tokens instead of model parameters, and during inference rather than training. By aggregating predictions from an ensemble of masked decoding contexts, we can robustly mitigate errors arising from visual token misinterpretations. Evaluations on benchmarks including CHAIR, THRONE, and MMBench demonstrate that DROPOUT DECODING significantly reduces object hallucinations (OH) and enhances both reliability and quality of LVLM outputs across diverse visual contexts. Code is released at https://github.com/kigb/DropoutDecoding . * Equal contribution. Work done during their research internship at Stony Brook University. † Joint last author. 3 We specifically refer to the tokens that are already in the input prompt to the text decoder. Concrete definition is in §3.1. 39th Conference on Neural Information Processing Systems (NeurIPS 2025).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3a373de7-a887-4b6d-9aa9-7d4ed569ab89Cited by top-tier papers3
- SCALE: Self-uncertainty Conditioned Adaptive Looking and Execution for Vision-Language-Action ModelsHyeonbeom Choi, Daechul Ahn, Youhan Lee, Taewook Kang et al.ICML 2026 · 2 citations
- DEGAP: Dynamic Entropy-Guided Attention Perturbation for Contrastive Decoding in Large Vision-Language ModelsHyein Seo, Yuna Jeong, Mingyu Kang, Junhyeong Park et al.ICML 2026
- Teaching VLMs to Admit Uncertainty in OCR from Lossy Visual InputsShuhao Guan, Moule Lin, Cheng Xu, Jinman Zhao et al.ICLR 2026
Builds on23
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- Bayesian Deep Learning and a Probabilistic Perspective of GeneralizationAndrew Gordon Wilson, Pavel IzmailovNeurIPS 2020 · 845 citations
- Ferret: Refer and Ground Anything Anywhere at Any GranularityHaoxuan You, Haotian Zhang, Zhe Gan, Xianzhi Du et al.ICLR 2024 · 515 citations
- Evaluating Object Hallucination in Large Vision-Language ModelsYifan Li, Yifan Du, Kun Zhou, Jinpeng Wang et al.EMNLP 2023 · 344 citations
Related papers
- On Epistemic Uncertainty of Visual Tokens for Object Hallucinations in Large Vision-Language ModelsHoigi Seo, Dong Un Kang, Hyunjin Cho, Joohoon Lee et al.NeurIPS 2025 · 4 citations
- Detecting Misbehaviors of Large Vision-Language Models by Evidential Uncertainty QuantificationTao Huang, Rui Wang, Xiaofei Liu, Yi Qin et al.ICLR 2026 · 4 citations
- Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language ModelsNanxing Hu, Xiaoyue Duan, Jinchao Zhang, Guoliang KangACM MM 2025 · 1 citation
- Revisit What You See: Revealing Visual Semantics in Vision Tokens to Guide LVLM DecodingBeomsik Cho, Jaehyung KimACL 2026
- Dropout Prompt Learning: Towards Robust and Adaptive Vision-Language ModelsBiao Chen, Lin Zuo, Mengmeng Jing, Kunbin He et al.AAAI 2026
