Effective Synthetic Data and Test-Time Adaptation for OCR Correction
Shuhao Guan, Cheng Xu, Moule Lin, Derek Greene
Abstract
Post-OCR technology is used to correct errors in the text produced by OCR systems. This study introduces a method for constructing post-OCR synthetic data with different noise levels using weak supervision. We define Character Error Rate (CER) thresholds for "effective" and "ineffective" synthetic data, allowing us to create more useful multi-noise level synthetic datasets. Furthermore, we propose Self-Correct-Noise Test-Time Adaptation (SCN-TTA), which combines self-correction and noise generation mechanisms. SCN-TTA allows a model to dynamically adjust to test data without relying on labels, effectively handling proper nouns in long texts and further reducing CER. In our experiments we evaluate a range of models, including multiple PLMs and LLMs. Results indicate that our method yields models that are effective across diverse text types. Notably, the ByT5 model achieves a CER reduction of 68.67% without relying on manually annotated data 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d523c73e-39aa-4186-be1b-50aaad5ce869Cited by top-tier papers4
- Table-Critic: A Multi-Agent Framework for Collaborative Criticism and Refinement in Table ReasoningPeiying Yu, Guoxin Chen, Jingjing WangACL 2025 · 30 citations
- LLaVA Steering: Visual Instruction Tuning with 500x Fewer Parameters through Modality Linear Representation-SteeringJinhe Bi, Yujun Wang, Haokun Chen, Xun Xiao et al.ACL 2025
- Teaching VLMs to Admit Uncertainty in OCR from Lossy Visual InputsShuhao Guan, Moule Lin, Cheng Xu, Jinman Zhao et al.ICLR 2026
- PreP-OCR: A Complete Pipeline for Document Image Restoration and Enhanced OCR AccuracyShuhao Guan, Moule Lin, Cheng Xu, Xinyi Liu et al.ACL 2025
Builds on7
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- Exploring and Predicting Transferability across NLP TasksTu Vu, Tong Wang, Tsendsuren Munkhdalai, Alessandro Sordoni et al.EMNLP 2020 · 104 citations
- Post-OCR Document Correction with Large Ensembles of Character Sequence-to-Sequence ModelsJuan Antonio Ramirez-Orta, Eduardo Xamena, Ana Gabriela Maguitman, Evangelos E. Milios et al.AAAI 2022 · 20 citations
- Beware of Model Collapse! Fast and Stable Test-time Adaptation for Robust Question AnsweringYi Su, Yixin Ji, Juntao Li, Hai Ye et al.EMNLP 2023 · 2 citations
Related papers
- Self-Supervised Text Erasing with Controllable Image SynthesisGangwei Jiang, Shiyao Wang, Tiezheng Ge, Yuning Jiang et al.ACM MM 2022 · 10 citations
- Optimized Tokenization for Transcribed Error CorrectionTomer Wullach, Shlomo E. ChazanEMNLP 2023 · 1 citation
- Back Translation for Speech-to-text Translation Without TranscriptsQingkai Fang, Yang FengACL 2023 · 9 citations
- OCR Post Correction for Endangered Language TextsShruti Rijhwani, Antonios Anastasopoulos, Graham NeubigEMNLP 2020 · 1 citation
- Once-More: Continuous Self-Correction for Large Language Models via Perplexity-Guided InterventionJiaxun Gao, Him Wai (Michael) Ng, Z. Jane WangICLR 2026
