Teaching VLMs to Admit Uncertainty in OCR from Lossy Visual Inputs
Shuhao Guan, Moule Lin, Cheng Xu, Jinman Zhao, Derek Greene
摘要
Vision-language models (VLMs) are increasingly replacing traditional OCR pipelines. However, they often hallucinate on lossy visual inputs, such as visually degraded document images, producing fluent yet incorrect text without signaling uncertainty. This occurs because current post-training emphasizes accuracy, which encourages models to guess even when uncertain. The problem persists in state-of-the-art systems and severely impacts OCR reliability. To improve the trustworthiness of OCR on degraded documents, we propose uncertainty-aware OCR. Rather than suppressing guesses, our model transcribes while explicitly bracketing spans it deems unreliable with uncertainty tags. To train our model, we use Group Relative Policy Optimization (GRPO). We define usage rules for uncertainty tags and an evaluation protocol, introducing a pseudo-labeled cold start and a multi-objective reward that balances transcription accuracy and uncertainty coverage while preventing reward hacking. We explore different combinations of cold-start and reward granularity. We also assess the effect of reward parameters in preventing reward hacking and improving the corresponding metrics. Furthermore, we introduce Blur-OCR, a challenging benchmark for uncertainty-aware OCR on degraded document images under lossy visual conditions. In extensive experiments, our model maintains transcription accuracy while achieving an uncertainty tag F1 score of 0.685. Data and code are available at https://github.com/NikoGuan/Uncertainty_OCR .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper18
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Real-Time Scene Text Detection with Differentiable BinarizationMinghui Liao, Zhaoyi Wan, Cong Yao, Kai Chen 等AAAI 2020 · 被引用 818 次
- TrOCR: Transformer-Based Optical Character Recognition with Pre-trained ModelsMinghao Li, Tengchao Lv, Jingye Chen, Lei Cui 等AAAI 2023 · 被引用 607 次
- Pix2Struct: Screenshot Parsing as Pretraining for Visual Language UnderstandingKenton Lee, Mandar Joshi, Iulia Raluca Turc, Hexiang Hu 等ICML 2023 · 被引用 426 次
相关 Paper
- Seeing is Believing? Mitigating OCR Hallucinations in Multimodal Large Language ModelsZhentao He, Can Zhang, Ziheng Wu, Zhenghao Chen 等NeurIPS 2025 · 被引用 14 次
- UCPO: Uncertainty-Aware Policy OptimizationXianzhou Zeng, Jing Huang, Chunmei Xie, Gongrui Nan 等ICML 2026
- From Exploration to Exploitation: A Two-Stage Entropy RLVR Approach for Noise-Tolerant MLLM TrainingDonglai Xu, Hongzheng Yang, Yuzhi Zhao, Pingping Zhang 等CVPR 2026 · 被引用 4 次
- Learning What to Trust: Bayesian Prior-Guided Optimization for Visual GenerationRuiying Liu, Yuanzhi Liang, Haibin Huang, Tianshu Yu 等CVPR 2026 · 被引用 4 次
- Calibration-Aware Policy Optimization for Reasoning LLMsZiqi Wang, Xingzhou Lou, Meiqi Wu, Zhengqi Wen 等ACL 2026 · 被引用 2 次
