Reading or Reasoning? Format Decoupled Reinforcement Learning for Document OCR
Yufeng Zhong, Lei Chen, Zhixiong Zeng, Xuanle Zhao, Deyang Jiang, Liming Zheng, Jing Huang, Haibo Qiu, Peng Shi, Siqi Yang, Lin Ma
摘要
Reading text from images or scanned documents via OCR models has been a longstanding focus of researchers. Intuitively, text reading is perceived as a straightforward perceptual task, and existing work primarily focuses on constructing enriched data engineering to enhance SFT capabilities. In this work, we observe that even advanced OCR models exhibit significantly higher entropy in formatted text (e.g., formula, table, etc.) compared to plain text, often by an order of magnitude. These statistical patterns reveal that advanced OCR models struggle with high output uncertainty when dealing with format sensitive document, suggesting that reasoning over diverse reading pathways may improve OCR performance. To address this, we propose format decoupled reinforcement learning (FD-RL), which leverages high-entropy patterns for targeted optimization. Our approach employs entropy-based data filtration strategy to identify format-intensive instances, and adopt format decoupled rewards tailored to different format types, enabling format-level validation rather than token-level memorization. FD-RL achieves an average score of 90.41 on OmniDocBench, setting a new record for end-to-end models on this highly popular benchmark. More importantly, we conduct comprehensive ablation studies over data, training, filtering, and rewarding strategies, thoroughly validating their effectiveness.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM ReasoningShenzhi Wang, Le Yu, Chang Gao, Chujie Zheng 等NeurIPS 2025 · 被引用 592 次
- Syntax-Aware Network for Handwritten Mathematical Expression RecognitionYe Yuan, Xiao Liu, Wondimu Dikubab, Hui Liu 等CVPR 2022 · 被引用 74 次
- OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive AnnotationsLinke Ouyang, Yuan Qu, Hongbin Zhou, Jiawei Zhu 等CVPR 2025
- POINTS-Reader: Distillation-Free Adaptation of Vision-Language Models for Document ConversionYuan Liu, Zhongyin Zhao, Le Tian, Haicheng Wang 等EMNLP 2025
相关 Paper
- Dissecting Post-Training: Uncovering the Complementary Roles of SFT and RL for Document ParsingJun-Peng Jiang, An-Yang Ji, Shiyin Lu, Guodong Zheng 等ICML 2026
- OvisOCR: End-to-End Document Parsing via Aligning Specialized Perception with General ReasoningJun-Peng Jiang, Shiyin Lu, An-Yang Ji, Yinglun Li 等ICML 2026
- TRIE: End-to-End Text Reading and Information Extraction for Document UnderstandingPeng Zhang, Yunlu Xu, Zhanzhan Cheng, Shiliang Pu 等ACM MM 2020 · 被引用 113 次
- Towards Unconstrained End-to-End Text SpottingSiyang Qin, Alessandro Bissacco, Michalis Raptis, Yasuhisa Fujii 等ICCV 2019 · 被引用 138 次
- UniDocVLM: Enhancing Visual Reasoning and Document Understanding for VLM via Reinforcement LearningZongsheng Cao, Anran Liu, Jun Xie, Lang Chen 等KDD 2026
