PreP-OCR: A Complete Pipeline for Document Image Restoration and Enhanced OCR Accuracy
Shuhao Guan, Moule Lin, Cheng Xu, Xinyi Liu, Jinman Zhao, Jiexin Fan, Qi Xu, Derek Greene
摘要
This paper introduces PreP-OCR, a two-stage pipeline that combines document image restoration with semantic-aware post-OCR correction to enhance both visual clarity and textual consistency, thereby improving text extraction from degraded historical documents. First, we synthesize document-image pairs from plaintext, rendering them with diverse fonts and layouts and then applying a randomly ordered set of degradation operations. An image restoration model is trained on this synthetic data, using multi-directional patch extraction and fusion to process large images. Second, a ByT5 post-OCR model, fine-tuned on synthetic historical text pairs, addresses remaining OCR errors. Detailed experiments on 13,831 pages of real historical documents in English, French, and Spanish show that the PreP-OCR pipeline reduces character error rates by 63.9-70.3% compared to OCR on raw images. Our pipeline demonstrates the potential of integrating image restoration with linguistic error correction for digitizing historical archives.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Table-Critic: A Multi-Agent Framework for Collaborative Criticism and Refinement in Table ReasoningPeiying Yu, Guoxin Chen, Jingjing WangACL 2025 · 被引用 30 次
- LLaVA Steering: Visual Instruction Tuning with 500x Fewer Parameters through Modality Linear Representation-SteeringJinhe Bi, Yujun Wang, Haokun Chen, Xun Xiao 等ACL 2025
- Teaching VLMs to Admit Uncertainty in OCR from Lossy Visual InputsShuhao Guan, Moule Lin, Cheng Xu, Jinman Zhao 等ICLR 2026
它引用的顶会 Paper16
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Restormer: Efficient Transformer for High-Resolution Image RestorationSyed Waqas Zamir, Aditya Arora, Salman Khan, Munawar Hayat 等CVPR 2022 · 被引用 3,348 次
- DeblurGAN-v2: Deblurring (Orders-of-Magnitude) Faster and BetterOrest Kupyn, Tetiana Martyniuk, Junru Wu, Zhangyang WangICCV 2019 · 被引用 1,100 次
- Rethinking Coarse-to-Fine Approach in Single Image DeblurringSung-Jin Cho, Seo-Won Ji, Jun-Pyo Hong, Seung-Won Jung 等ICCV 2021 · 被引用 799 次
- ResShift: Efficient Diffusion Model for Image Super-resolution by Residual ShiftingZongsheng Yue, Jianyi Wang, Chen Change LoyNeurIPS 2023 · 被引用 646 次
相关 Paper
- Reviving Cultural Heritage: A Novel Approach for Comprehensive Historical Document RestorationYuyi Zhang, Peirong Zhang, Zhenhua Yang, Pengyu Yan 等ACL 2025 · 被引用 5 次
- Predicting the Original Appearance of Damaged Historical DocumentsZhenhua Yang, Dezhi Peng, Yongxin Shi, Yuyi Zhang 等AAAI 2025 · 被引用 8 次
- PHD: Pixel-Based Language Modeling of Historical DocumentsNadav Borenstein, Phillip Rust, Desmond Elliott, Isabelle AugensteinEMNLP 2023
- Draft, Verify, Restore: Self-Refining Historical Inscription Restoration with a Unified MLLMYuyi Zhang, Junle Liu, Peirong Zhang, Jianliang Liu 等ACL 2026
- Post-OCR Document Correction with Large Ensembles of Character Sequence-to-Sequence ModelsJuan Antonio Ramirez-Orta, Eduardo Xamena, Ana Gabriela Maguitman, Evangelos E. Milios 等AAAI 2022 · 被引用 20 次
