Post-OCR Document Correction with Large Ensembles of Character Sequence-to-Sequence Models
Juan Antonio Ramirez-Orta, Eduardo Xamena, Ana Gabriela Maguitman, Evangelos E. Milios, Axel J. Soto
摘要
In this paper, we propose a novel method to extend sequence-to-sequence models to accurately process sequences much longer than the ones used during training while being sample- and resource-efficient, supported by thorough experimentation. To investigate the effectiveness of our method, we apply it to the task of correcting documents already processed with Optical Character Recognition (OCR) systems using sequence-to-sequence models based on characters. We test our method on nine languages of the ICDAR 2019 competition on post-OCR text correction and achieve a new state-of-the-art performance in five of them. The strategy with the best performance involves splitting the input document in character n-grams and combining their individual corrections into the final output using a voting scheme that is equivalent to an ensemble of a large number of sequence models. We further investigate how to weigh the contributions from each one of the members of this ensemble. Our code for post-OCR correction is shared at https://github.com/jarobyte91/post_ocr_correction.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper1
相关 Paper
- Copy That! Editing Sequences by Copying SpansSheena Panthaplackel, Miltiadis Allamanis, Marc BrockschmidtAAAI 2021 · 被引用 28 次
- PreP-OCR: A Complete Pipeline for Document Image Restoration and Enhanced OCR AccuracyShuhao Guan, Moule Lin, Cheng Xu, Xinyi Liu 等ACL 2025
- Efficient OCR for Building a Diverse Digital HistoryJacob Carlson, Tom Bryan, Melissa DellACL 2024 · 被引用 6 次
- Improving Code Extraction from Coding Screencasts Using a Code-Aware Encoder-Decoder ModelAbdulkarim Malkadi, Ahmad Tayeb, Sonia HaiducASE 2023 · 被引用 5 次
- Sequence-to-Action: Grammatical Error Correction with Action Guided Sequence GenerationJiquan Li, Junliang Guo, Yongxin Zhu, Xin Sheng 等AAAI 2022 · 被引用 29 次
