When Diffusion Language Models Hesitate: Detecting and Correcting Visual Hallucinations via Confidence Fluctuation
Wenzheng Song, Pei Chen, Yichen Tan, Zejian Li, Lingyun Sun
摘要
Multi-modal Diffusion Language Models (MDLMs) have emerged as a powerful alternative to autoregressive models in vision understanding, offering advantages in bidirectional context modeling and parallel decoding. However, existing MDLMs suffer from visual hallucinations due to the static nature of visual perception.
Unlike autoregressive models, MDLMs lack the sequential dependency to dynamically interact with visual content. Therefore, MDLMs rely on fixed visual features encoded at initialization, causing the denoising process to drift toward language priors and lose its anchor to visual evidence. In this paper, we propose VGR (Visual-Guided Refinement), a framework that enables MDLMs to revisit visual details by exploiting diffusion dynamics. Our key insight is that the temporal trajectory of confidence during denoising reveals intrinsic uncertainty: while grounded tokens converge smoothly, hallucinated ones exhibit pronounced confidence fluctuation. VGR utilizes this fluctuation signal to detect uncertain spans and corrects them through targeted visual evidence extraction and in-place remasking. Extensive experiments on image captioning and hallucination evaluation benchmarks demonstrate that our method reduces hallucinations and recalls more details. We release our code at https: //github.com/SongWZ3214/VGR.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper21
- Large Language Diffusion ModelsShen Nie, Fengqi Zhu, Zebin You, Xiaolu Zhang 等NeurIPS 2025 · 被引用 949 次
- Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language ModelsYushi Hu, Weijia Shi, Xingyu Fu, Dan Roth 等NeurIPS 2024 · 被引用 373 次
- DeepEyes: Incentivizing "Thinking with Images" via Reinforcement LearningZiwei Zheng, Michael Yang, Jack Hong, Chenxiao Zhao 等ICLR 2026 · 被引用 321 次
- Remasking Discrete Diffusion Models with Inference-Time ScalingGuanghan Wang, Yair Schiff, Subham S. Sahoo, Volodymyr KuleshovNeurIPS 2025 · 被引用 199 次
- LLaDA-V: Large Language Diffusion Models with Visual Instruction TuningZebin You, Shen Nie, Xiaolu Zhang, JUN ZHOU 等CVPR 2026 · 被引用 154 次
相关 Paper
- Hallucination Begins Where Saliency DropsXiaofeng Zhang, Yuanchao Zhu, Chaochen Gu, Xiaosong Yuan 等ICLR 2026 · 被引用 11 次
- VCGD: Visual Clue Guided Decoding with Caption Model for Mitigating Hallucination in Multimodal Large Language ModelsGuoqing Chen, Fu Zhang, Bingqian Liu, Chenglong Lu 等AAAI 2026
- Look Twice Before You Answer: Memory-Space Visual Retracing for Hallucination Mitigation in Multimodal Large Language ModelsXin Zou, Yizhou Wang, Yibo Yan, Yuanhuiyi Lyu 等ICML 2025 · 被引用 1 次
- Tell Model Where to Look: Mitigating Hallucinations in MLLMs by Vision-Guided AttentionJianfei Zhao, Feng Zhang, Xin Sun, Chong Feng 等CVPR 2026 · 被引用 6 次
- Generative Multimodal Pretraining with Discrete Diffusion Timestep TokensKaihang Pan, Wang Lin, Zhongqi Yue, Tenglong Ao 等CVPR 2025
