On the Generalization of Handwritten Text Recognition Models
Carlos Garrido-Munoz, Jorge Calvo-Zaragoza
Abstract
Recent advances in Handwritten Text Recognition (HTR) have led to significant reductions in transcription errors on standard benchmarks under the i.i.d. assumption, thus focusing on minimizing in-distribution (ID) errors. However, this assumption does not hold in real-world applications, which has motivated HTR research to explore Transfer Learning and Domain Adaptation techniques. In this work, we investigate the unaddressed limitations of HTR models in generalizing to out-of-distribution (OOD) data. We adopt the challenging setting of Domain Generalization, where models are expected to generalize to OOD data without any prior access. To this end, we analyze 336 OOD cases from eight state-of-the-art HTR models across seven widely used datasets, spanning five languages. Additionally, we study how HTR models leverage synthetic data to generalize. We reveal that the most significant factor for generalization lies in the textual divergence between domains, followed by visual divergence. We demonstrate that the error of HTR models in OOD scenarios can be reliably estimated, with discrepancies falling below 10 points in 70% of cases. We identify the underlying limitations of HTR models, laying the foundation for future research to address this challenge. Code is available at github.com/carlos10garrido/HTR-OOD.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bbbe6301-9c12-4eef-9a63-b96c66d3538dBuilds on16
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- In Search of Lost Domain GeneralizationIshaan Gulrajani, David Lopez-PazICLR 2021 · 1,416 citations
- TrOCR: Transformer-Based Optical Character Recognition with Pre-trained ModelsMinghao Li, Tengchao Lv, Jingye Chen, Lei Cui et al.AAAI 2023 · 607 citations
- Learning to Diversify for Single Domain GeneralizationZijian Wang, Yadan Luo, Ruihong Qiu, Zi Huang et al.ICCV 2021 · 339 citations
Related papers
- JokerGAN: Memory-Efficient Model for Handwritten Text Generation with Text Line AwarenessJan Zdenek, Hideki NakayamaACM MM 2021 · 22 citations
- Digitizing Nepal's Written Heritage: A Comprehensive HTR Pipeline for Old Nepali ManuscriptsAnjali Sarawgi, Esteban Garces Arias, Christof ZotterACL 2026 · 2 citations
- Automatic Transcription of Handwritten Old Occitan LanguageEsteban Garces Arias, Vallari Pai, Matthias Schöffel, Christian Heumann et al.EMNLP 2023 · 2 citations
- Semantic-Discriminative Mixup for Generalizable Sensor-based Cross-domain Activity RecognitionWang Lu, Jindong Wang, Yiqiang Chen, Sinno Jialin Pan et al.UbiComp 2022 · 61 citations
- What if We Only Use Real Datasets for Scene Text Recognition? Toward Scene Text Recognition With Fewer LabelsJeonghun Baek, Yusuke Matsui, Kiyoharu AizawaCVPR 2021
