Ground What You See: Hallucination-Resistant MLLMs via Caption Feedback, Diversity-Aware Sampling, and Conflict Regularization
Miao Pan, Wangjie Gan, Jintao Chen, Wenqi Zhang, Bing Sun, Jianwei Yin, Xuhong Zhang
摘要
Multimodal large language models (MLLMs) have achieved significant results in various tasks, but their practical application is still severely constrained by hallucination issues, which are particularly prominent in reinforcement learning (RL) optimization processes. This paper systematically analyzes the causes of hallucinations in MLLM under RL training, identifying three key factors: (1) The model relies heavily on chained visual reasoning to guide decision-making during RL training. Thus, error and irrelevant information in visual reasoning can easily cause hallucinations, including inaccurate initial visual descriptions that anchor subsequent inferences to incorrect information, as well as redundant and broad inferential information; (2) Insufficient exploration diversity during the policy optimization phase, causing the model to output overly confident results; (3) The destructive conflict between different samples during optimization is a key factor that leads to false associations and unstable parameter updates. To address these issues, we propose a solution framework comprising three core modules. First, to improve the accuracy of visual localization, we add planning and caption stages before thinking and answer stages. To enhance initial visual descriptions ability, we allow LLMs to respond based solely on the caption and provide corresponding caption reward based on the quality of the response. Second, to enhance exploration capabilities, we classify samples based on the mean and variance of the reward distribution and select samples with high reward variance for training, thereby increasing the model's focus on diverse samples. Finally, to mitigate conflicts between training samples, we identify neural tangent kernel (NTK) similarity as the key factor. Rather than minimizing it uniformly, we regulate NTK similarity by grouping sample pairs based on a similarity threshold. An In-foNCE loss is then applied to pull dissimilar pairs closer and push overly similar ones apart, guiding interactions toward a balanced range. The experimental results demonstrate that the proposed method significantly reduces the hallucination rate and effectively improves the inference accuracy of MLLMs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual ContextsPan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu 等ICLR 2024 · 被引用 1,472 次
- Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language ModelsWenxuan Huang, Bohan Jia, Shaosheng Cao, Zheyu Ye 等ICLR 2026 · 被引用 670 次
- Video-R1: Reinforcing Video Reasoning in MLLMsKaituo Feng, Kaixiong Gong, Bohao Li, Zonghao Guo 等NeurIPS 2025 · 被引用 528 次
- Evaluating Object Hallucination in Large Vision-Language ModelsYifan Li, Yifan Du, Kun Zhou, Jinpeng Wang 等EMNLP 2023 · 被引用 344 次
- Q-Insight: Understanding Image Quality via Visual Reinforcement LearningWeiqi Li, Xuanyu Zhang, Shijie Zhao, Yabin Zhang 等NeurIPS 2025 · 被引用 117 次
相关 Paper
- VCGD: Visual Clue Guided Decoding with Caption Model for Mitigating Hallucination in Multimodal Large Language ModelsGuoqing Chen, Fu Zhang, Bingqian Liu, Chenglong Lu 等AAAI 2026
- Robust Multimodal Large Language Models Against Modality ConflictZongmeng Zhang, Wengang Zhou, Jie Zhao, Houqiang LiICML 2025
- Understanding the Role of Hallucination in Reinforcement Post-Training of Multimodal Reasoning ModelsGengwei Zhang, Jie Peng, Zhen Tan, Mufan Qiu 等CVPR 2026 · 被引用 1 次
- InEx: Hallucination Mitigation via Introspection and Cross-Modal Multi-Agent CollaborationZhongyu Yang, Yingfang Yuan, Xuanming Jiang, Baoyi An 等AAAI 2026 · 被引用 5 次
- Visual Attention Reasoning via Hierarchical Search and Self-VerificationWei Cai, Jian Zhao, Yuchen Yuan, Tianle Zhang 等ACL 2026
