Seeing is Believing? Mitigating OCR Hallucinations in Multimodal Large Language Models
Zhentao He, Can Zhang, Ziheng Wu, Zhenghao Chen, Yufei Zhan, Yifan Li, Zhao Zhang, Xian Wang, Minghui Qiu
Abstract
Recent advancements in multimodal large language models (MLLMs) have enhanced document understanding by integrating textual and visual information. However, existing models exhibit incompleteness within their paradigm in real-world scenarios, particularly under visual degradation (e.g., blur, occlusion, low contrast). In such conditions, the current response paradigm often fails to adequately perceive visual degradation and ambiguity, leading to overreliance on linguistic priors or misaligned visual-textual reasoning. This difficulty in recognizing uncertainty frequently results in the generation of hallucinatory content, especially when a precise answer is not feasible. To better demonstrate and analyze this phenomenon and problem, we propose KIE-HVQA, the first benchmark dedicated to evaluating OCR hallucination in degraded document understanding. This dataset includes test samples spanning identity cards, invoices, and prescriptions, with simulated real-world degradations and pixel-level annotations for OCR reliability. This setup allows for evaluating models' capacity, under degraded input, to distinguish reliable visual information and answer accordingly, thereby highlighting the challenge of avoiding hallucination on uncertain data. To achieve vision-faithful reasoning and thereby avoid the aforementioned issues, we further introduce a Group Relative Policy Optimization (GRPO)-based framework featuring a novel reward mechanism. By incorporating a self-awareness of visual uncertainty and an analysis method that initiates refusal to answer to increase task difficulty within our supervised fine-tuning and reinforcement learning framework, we successfully mitigated hallucinations in ambiguous regions. Experiments on Qwen2.5-VL demonstrate that our 7B-parameter model achieves a ∼28% absolute improvement in hallucination-free accuracy over GPT-4o on KIE-HVQA and there is no significant performance drop in standard tasks, highlighting both effectiveness and robustness. This work advances the development of reliable MLLMs for real-world document analysis by addressing critical challenges in visual-linguistic alignment under degradation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b2961ee5-981d-47fb-9de3-0029acf1c06eCited by top-tier papers5
- TextPecker: Rewarding Structural Anomaly Quantification for Enhancing Visual Text RenderingHanshen Zhu, Yuliang Liu, Xuecheng Wu, An-Lan Wang et al.CVPR 2026 · 15 citations
- Detached Skip-Links and -Probe: Decoupling Feature Aggregation from Gradient Propagation for MLLM OCRZiye Yuan, Ruchang Yao, Chengxin Zheng, Yusheng Zhao et al.ICML 2026
- AD-BTS: Adaptive Dual-Branch Token Sparsification via Spatial Information DensityXinpei Gao, Xin Luo, Ming Liu, Chunjiang Wang et al.ICML 2026
- Teaching VLMs to Admit Uncertainty in OCR from Lossy Visual InputsShuhao Guan, Moule Lin, Cheng Xu, Jinman Zhao et al.ICLR 2026
- MessToClean: Evidence-Grounded Structure-Preserving Reconstruction for Real-World Degraded Exam Paper ImagesJiayi Tuo, Cheng Tang, Zihan Wang, Chenyue Zhou et al.ACL 2026
Builds on11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Visual-RFT: Visual Reinforcement Fine-TuningZiyu Liu, Zeyi Sun, Yuhang Zang, Xiaoyi Dong et al.ICCV 2025 · 563 citations
- Reason-RFT: Reinforcement Fine-Tuning for Visual Reasoning of Vision Language ModelsHuajie Tan, Yuheng Ji, Xiaoshuai Hao, Xiansheng Chen et al.NeurIPS 2025 · 45 citations
- LayoutLLM: Layout Instruction Tuning with Large Language Models for Document UnderstandingChuwei Luo, Yufan Shen, Zhaoqing Zhu, Qi Zheng et al.CVPR 2024 · 39 citations
Related papers
- Hallu-PI: Evaluating Hallucination in Multi-modal Large Language Models within Perturbed InputsPeng Ding, Jingyu Wu, Jun Kuang, Dan Ma et al.ACM MM 2024 · 8 citations
- From Exploration to Exploitation: A Two-Stage Entropy RLVR Approach for Noise-Tolerant MLLM TrainingDonglai Xu, Hongzheng Yang, Yuzhi Zhao, Pingping Zhang et al.CVPR 2026 · 4 citations
- AD-FM: Multimodal LLMs for Anomaly Detection via Multi-Stage Reasoning and Fine-Grained Reward OptimizationJingyi Liao, Yongyi Su, Rong-Cheng Tu, Zhao Jin et al.AAAI 2026
- ELV-Halluc: Benchmarking Semantic Aggregation Hallucinations in Video UnderstandingHao Lu, Jiahao Wang, Yaolun Zhang, Ruohui Wang et al.CVPR 2026
- From Detection to Diagnosis: Advancing Hallucination Analysis with Automated Data SynthesisYanyi Liu, Qingwen Yang, Tiezheng Guo, Feiyu Qu et al.AAAI 2026
