PreP-OCR: A Complete Pipeline for Document Image Restoration and Enhanced OCR Accuracy
Shuhao Guan, Moule Lin, Cheng Xu, Xinyi Liu, Jinman Zhao, Jiexin Fan, Qi Xu, Derek Greene
Abstract
This paper introduces PreP-OCR, a two-stage pipeline that combines document image restoration with semantic-aware post-OCR correction to enhance both visual clarity and textual consistency, thereby improving text extraction from degraded historical documents. First, we synthesize document-image pairs from plaintext, rendering them with diverse fonts and layouts and then applying a randomly ordered set of degradation operations. An image restoration model is trained on this synthetic data, using multi-directional patch extraction and fusion to process large images. Second, a ByT5 post-OCR model, fine-tuned on synthetic historical text pairs, addresses remaining OCR errors. Detailed experiments on 13,831 pages of real historical documents in English, French, and Spanish show that the PreP-OCR pipeline reduces character error rates by 63.9-70.3% compared to OCR on raw images. Our pipeline demonstrates the potential of integrating image restoration with linguistic error correction for digitizing historical archives.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b8665d90-1534-4887-8ded-5b3dce7cb004Cited by top-tier papers3
- Table-Critic: A Multi-Agent Framework for Collaborative Criticism and Refinement in Table ReasoningPeiying Yu, Guoxin Chen, Jingjing WangACL 2025 · 30 citations
- LLaVA Steering: Visual Instruction Tuning with 500x Fewer Parameters through Modality Linear Representation-SteeringJinhe Bi, Yujun Wang, Haokun Chen, Xun Xiao et al.ACL 2025
- Teaching VLMs to Admit Uncertainty in OCR from Lossy Visual InputsShuhao Guan, Moule Lin, Cheng Xu, Jinman Zhao et al.ICLR 2026
Builds on16
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Restormer: Efficient Transformer for High-Resolution Image RestorationSyed Waqas Zamir, Aditya Arora, Salman Khan, Munawar Hayat et al.CVPR 2022 · 3,348 citations
- DeblurGAN-v2: Deblurring (Orders-of-Magnitude) Faster and BetterOrest Kupyn, Tetiana Martyniuk, Junru Wu, Zhangyang WangICCV 2019 · 1,100 citations
- Rethinking Coarse-to-Fine Approach in Single Image DeblurringSung-Jin Cho, Seo-Won Ji, Jun-Pyo Hong, Seung-Won Jung et al.ICCV 2021 · 799 citations
- ResShift: Efficient Diffusion Model for Image Super-resolution by Residual ShiftingZongsheng Yue, Jianyi Wang, Chen Change LoyNeurIPS 2023 · 646 citations
Related papers
- Reviving Cultural Heritage: A Novel Approach for Comprehensive Historical Document RestorationYuyi Zhang, Peirong Zhang, Zhenhua Yang, Pengyu Yan et al.ACL 2025 · 5 citations
- Predicting the Original Appearance of Damaged Historical DocumentsZhenhua Yang, Dezhi Peng, Yongxin Shi, Yuyi Zhang et al.AAAI 2025 · 8 citations
- PHD: Pixel-Based Language Modeling of Historical DocumentsNadav Borenstein, Phillip Rust, Desmond Elliott, Isabelle AugensteinEMNLP 2023
- Draft, Verify, Restore: Self-Refining Historical Inscription Restoration with a Unified MLLMYuyi Zhang, Junle Liu, Peirong Zhang, Jianliang Liu et al.ACL 2026
- Post-OCR Document Correction with Large Ensembles of Character Sequence-to-Sequence ModelsJuan Antonio Ramirez-Orta, Eduardo Xamena, Ana Gabriela Maguitman, Evangelos E. Milios et al.AAAI 2022 · 20 citations
