Teaching VLMs to Admit Uncertainty in OCR from Lossy Visual Inputs
Shuhao Guan, Moule Lin, Cheng Xu, Jinman Zhao, Derek Greene
Abstract
Vision-language models (VLMs) are increasingly replacing traditional OCR pipelines. However, they often hallucinate on lossy visual inputs, such as visually degraded document images, producing fluent yet incorrect text without signaling uncertainty. This occurs because current post-training emphasizes accuracy, which encourages models to guess even when uncertain. The problem persists in state-of-the-art systems and severely impacts OCR reliability. To improve the trustworthiness of OCR on degraded documents, we propose uncertainty-aware OCR. Rather than suppressing guesses, our model transcribes while explicitly bracketing spans it deems unreliable with uncertainty tags. To train our model, we use Group Relative Policy Optimization (GRPO). We define usage rules for uncertainty tags and an evaluation protocol, introducing a pseudo-labeled cold start and a multi-objective reward that balances transcription accuracy and uncertainty coverage while preventing reward hacking. We explore different combinations of cold-start and reward granularity. We also assess the effect of reward parameters in preventing reward hacking and improving the corresponding metrics. Furthermore, we introduce Blur-OCR, a challenging benchmark for uncertainty-aware OCR on degraded document images under lossy visual conditions. In extensive experiments, our model maintains transcription accuracy while achieving an uncertainty tag F1 score of 0.685. Data and code are available at https://github.com/NikoGuan/Uncertainty_OCR .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b5c90534-b815-44ac-a1c7-8ea6b59bfe60Cited by top-tier papers1
Ask how each one uses itBuilds on18
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Real-Time Scene Text Detection with Differentiable BinarizationMinghui Liao, Zhaoyi Wan, Cong Yao, Kai Chen et al.AAAI 2020 · 818 citations
- TrOCR: Transformer-Based Optical Character Recognition with Pre-trained ModelsMinghao Li, Tengchao Lv, Jingye Chen, Lei Cui et al.AAAI 2023 · 607 citations
- Pix2Struct: Screenshot Parsing as Pretraining for Visual Language UnderstandingKenton Lee, Mandar Joshi, Iulia Raluca Turc, Hexiang Hu et al.ICML 2023 · 426 citations
Related papers
- Seeing is Believing? Mitigating OCR Hallucinations in Multimodal Large Language ModelsZhentao He, Can Zhang, Ziheng Wu, Zhenghao Chen et al.NeurIPS 2025 · 14 citations
- UCPO: Uncertainty-Aware Policy OptimizationXianzhou Zeng, Jing Huang, Chunmei Xie, Gongrui Nan et al.ICML 2026
- From Exploration to Exploitation: A Two-Stage Entropy RLVR Approach for Noise-Tolerant MLLM TrainingDonglai Xu, Hongzheng Yang, Yuzhi Zhao, Pingping Zhang et al.CVPR 2026 · 4 citations
- Learning What to Trust: Bayesian Prior-Guided Optimization for Visual GenerationRuiying Liu, Yuanzhi Liang, Haibin Huang, Tianshu Yu et al.CVPR 2026 · 4 citations
- Calibration-Aware Policy Optimization for Reasoning LLMsZiqi Wang, Xingzhou Lou, Meiqi Wu, Zhengqi Wen et al.ACL 2026 · 2 citations
