Effective Synthetic Data and Test-Time Adaptation for OCR Correction
Shuhao Guan, Cheng Xu, Moule Lin, Derek Greene
摘要
Post-OCR technology is used to correct errors in the text produced by OCR systems. This study introduces a method for constructing post-OCR synthetic data with different noise levels using weak supervision. We define Character Error Rate (CER) thresholds for "effective" and "ineffective" synthetic data, allowing us to create more useful multi-noise level synthetic datasets. Furthermore, we propose Self-Correct-Noise Test-Time Adaptation (SCN-TTA), which combines self-correction and noise generation mechanisms. SCN-TTA allows a model to dynamically adjust to test data without relying on labels, effectively handling proper nouns in long texts and further reducing CER. In our experiments we evaluate a range of models, including multiple PLMs and LLMs. Results indicate that our method yields models that are effective across diverse text types. Notably, the ByT5 model achieves a CER reduction of 68.67% without relying on manually annotated data 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Table-Critic: A Multi-Agent Framework for Collaborative Criticism and Refinement in Table ReasoningPeiying Yu, Guoxin Chen, Jingjing WangACL 2025 · 被引用 30 次
- LLaVA Steering: Visual Instruction Tuning with 500x Fewer Parameters through Modality Linear Representation-SteeringJinhe Bi, Yujun Wang, Haokun Chen, Xun Xiao 等ACL 2025
- Teaching VLMs to Admit Uncertainty in OCR from Lossy Visual InputsShuhao Guan, Moule Lin, Cheng Xu, Jinman Zhao 等ICLR 2026
- PreP-OCR: A Complete Pipeline for Document Image Restoration and Enhanced OCR AccuracyShuhao Guan, Moule Lin, Cheng Xu, Xinyi Liu 等ACL 2025
它引用的顶会 Paper7
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman 等ICML 2023 · 被引用 6,966 次
- Exploring and Predicting Transferability across NLP TasksTu Vu, Tong Wang, Tsendsuren Munkhdalai, Alessandro Sordoni 等EMNLP 2020 · 被引用 104 次
- Post-OCR Document Correction with Large Ensembles of Character Sequence-to-Sequence ModelsJuan Antonio Ramirez-Orta, Eduardo Xamena, Ana Gabriela Maguitman, Evangelos E. Milios 等AAAI 2022 · 被引用 20 次
- Beware of Model Collapse! Fast and Stable Test-time Adaptation for Robust Question AnsweringYi Su, Yixin Ji, Juntao Li, Hai Ye 等EMNLP 2023 · 被引用 2 次
相关 Paper
- Self-Supervised Text Erasing with Controllable Image SynthesisGangwei Jiang, Shiyao Wang, Tiezheng Ge, Yuning Jiang 等ACM MM 2022 · 被引用 10 次
- Optimized Tokenization for Transcribed Error CorrectionTomer Wullach, Shlomo E. ChazanEMNLP 2023 · 被引用 1 次
- Back Translation for Speech-to-text Translation Without TranscriptsQingkai Fang, Yang FengACL 2023 · 被引用 9 次
- OCR Post Correction for Endangered Language TextsShruti Rijhwani, Antonios Anastasopoulos, Graham NeubigEMNLP 2020 · 被引用 1 次
- Once-More: Continuous Self-Correction for Large Language Models via Perplexity-Guided InterventionJiaxun Gao, Him Wai (Michael) Ng, Z. Jane WangICLR 2026
