TexOCR: Advancing Document OCR Models for Compilable Page-to-LaTeX Reconstruction
Chengye Wang, Lin Fu, Zexi Kuang, Yilun Zhao
摘要
Existing document OCR largely targets plain text or Markdown, discarding the structural and executable properties that make LaTeX essential for scientific publishing. We study page-level reconstruction of scientific PDFs into compilable LaTeX and introduce TEX-OCR-Bench, a benchmark, and TEXOCR-Train, a large-scale training corpus, for this task. TEXOCR-Bench features a multi-dimensional evaluation suite that jointly assesses transcription fidelity, structural faithfulness, and endto-end compilability. Leveraging TEXOCR-Train, we train a 2B-parameter model, TEX-OCR, using supervised fine-tuning (SFT) and reinforcement learning (RL) with verifiable rewards derived from LaTeX unit tests that directly enforce compilability and referential integrity. Experiments across 21 frontier models on TEXOCR-Bench show that existing systems frequently violate key document invariants, including consistent section structure, correct float placement, and valid label-reference links, which undermines compilation reliability and downstream usability. Our analysis further reveals that RL with verifiable rewards yields consistent improvements over SFT alone, particularly on structural and compilation metrics.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- S2ORC: The Semantic Scholar Open Research CorpusKyle Lo, Lucy Lu Wang, Mark Neumann, Rodney Kinney 等ACL 2020 · 被引用 424 次
- Nougat: Neural Optical Understanding for Academic DocumentsLukas Blecher, Guillem Cucurull, Thomas Scialom, Robert StojnicICLR 2024 · 被引用 243 次
- Table2LaTeX-RL: High-Fidelity LaTeX Code Generation from Table Images via Reinforced Multimodal Language ModelsJun Ling, Yao Qi, Tao Huang, Shibo Zhou 等NeurIPS 2025 · 被引用 9 次
- SmolDocling: An Ultra-Compact Vision-Language Model for End-To-End Multi-Modal Document ConversionAhmed S. Nassar, Matteo Omenetti, Maksym Lysak, Nikolaos Livathinos 等ICCV 2025 · 被引用 9 次
相关 Paper
- VerifyBench: A Systematic Benchmark for Evaluating Reasoning Verifiers Across DomainsXuzhao Li, Xuchen Li, Shiyu Hu, Yongzhen Guo 等AAAI 2026 · 被引用 16 次
- LaTeX2Layout: High-Fidelity, Scalable Document Layout Annotation Pipeline for Layout DetectionFeijiang Han, Zelong Wang, Bowen Wang, Xinxin Liu 等AAAI 2026 · 被引用 4 次
- ReFF: Reinforcing Format Faithfulness in Language Models Across Varied TasksJiashu Yao, Heyan Huang, Zeming Liu, Haoyu Wen 等AAAI 2025 · 被引用 1 次
- OCR Hinders RAG: Evaluating the Cascading Impact of OCR on Retrieval-Augmented GenerationJunyuan Zhang, Qintong Zhang, Bin Wang, Linke Ouyang 等ICCV 2025 · 被引用 10 次
- CSPO: Alleviating Reward Ambiguity for Structured Table-to-LaTeX GenerationYunfan Yang, Cuiling Lan, Jitao Sang, Yan LuACL 2026
