Automatic Transcription of Handwritten Old Occitan Language
Esteban Garces Arias, Vallari Pai, Matthias Schöffel, Christian Heumann, Matthias Aßenmacher
摘要
While existing neural network-based approaches have shown promising results in Handwritten Text Recognition (HTR) for highresource languages and standardized/machinewritten text, their application to low-resource languages often presents challenges, resulting in reduced effectiveness. In this paper, we propose an innovative HTR approach that leverages the Transformer architecture for recognizing handwritten Old Occitan language. Given the limited availability of data, which comprises only word pairs of graphical variants and lemmas, we develop and rely on elaborate data augmentation techniques for both text and image data. Our model combines a custom-trained Swin image encoder with a BERT text decoder, which we pre-train using a large-scale augmented synthetic data set and fine-tune on the small human-labeled data set. Experimental results reveal that our approach surpasses the performance of current state-ofthe-art models for Old Occitan HTR, including open-source Transformer-based models such as a fine-tuned TrOCR and commercial applications like Google Cloud Vision. To nurture further research and development, we make our models, data sets, and code publicly available: https://huggingface.co/misoda
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper4
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 被引用 3,632 次
- TrOCR: Transformer-Based Optical Character Recognition with Pre-trained ModelsMinghao Li, Tengchao Lv, Jingye Chen, Lei Cui 等AAAI 2023 · 被引用 607 次
相关 Paper
- JokerGAN: Memory-Efficient Model for Handwritten Text Generation with Text Line AwarenessJan Zdenek, Hideki NakayamaACM MM 2021 · 被引用 22 次
- ScrabbleGAN: Semi-Supervised Varying Length Handwritten Text GenerationSharon Fogel, Hadar Averbuch-Elor, Sarel Cohen, Shai Mazor 等CVPR 2020
- Learn to Augment: Joint Data Augmentation and Network Optimization for Text RecognitionCanjie Luo, Yuanzhi Zhu, Lianwen Jin, Yongpan WangCVPR 2020
- Text-DIAE: A Self-Supervised Degradation Invariant Autoencoder for Text Recognition and Document EnhancementMohamed Ali Souibgui, Sanket Biswas, Andrés Mafla, Ali Furkan Biten 等AAAI 2023 · 被引用 31 次
- KnowDA: All-in-One Knowledge Mixture Model for Data Augmentation in Low-Resource NLPYufei Wang, Jiayi Zheng, Can Xu, Xiubo Geng 等ICLR 2023 · 被引用 2 次
