Digitizing Nepal's Written Heritage: A Comprehensive HTR Pipeline for Old Nepali Manuscripts
Anjali Sarawgi, Esteban Garces Arias, Christof Zotter
Abstract
This paper presents the first end-to-end pipeline for Handwritten Text Recognition (HTR) for Old Nepali, a historically significant but lowresource language. We adopt a line-level transcription approach and systematically explore encoder-decoder architectures and data-centric techniques to improve recognition accuracy. Our best model achieves a Character Error Rate (CER) of 4.9%. In addition, we implement and evaluate decoding strategies and analyze tokenlevel confusions to better understand model behavior and error patterns. Although the evaluation dataset is confidential, we release our training code, model configurations, and evaluation scripts to support further research on HTR for low-resource historical scripts. However, the historical manuscripts considered in this study present several challenges, including diverse handwriting styles, degraded document quality, intricate conjunct forms, and limited annotated data (Nockels et al., 2024) . These constraints necessitate approaches that leverage transfer learning, data augmentation, and carefully designed model architectures (Garces Arias et al., 2023). In this paper, we address these challenges by conducting a comprehensive exploration of modern HTR techniques for Old Nepali manuscripts. We investigate transfer learning strategies to enable effective learning in low-resource settings, implementing a three-stage approach that adapts models from high-resource settings to target-domain data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5d410fb6-4ab3-4f2a-9425-97ac36d87b59Builds on6
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- TrOCR: Transformer-Based Optical Character Recognition with Pre-trained ModelsMinghao Li, Tengchao Lv, Jingye Chen, Lei Cui et al.AAAI 2023 · 607 citations
- A Contrastive Framework for Neural Text GenerationYixuan Su, Tian Lan, Yan Wang, Dani Yogatama et al.NeurIPS 2022 · 349 citations
- Min-k Sampling: Decoupling Truncation from Temperature Scaling via Relative Logit DynamicsYuanhao Ding, Meimingwei Li, Esteban Garces Arias, Matthias Aßenmacher et al.ACL 2026 · 4 citations
Related papers
- Automatic Transcription of Handwritten Old Occitan LanguageEsteban Garces Arias, Vallari Pai, Matthias Schöffel, Christian Heumann et al.EMNLP 2023 · 2 citations
- General Detection-based Text Line RecognitionRaphaël Baena, Syrine Kalleli, Mathieu AubryNeurIPS 2024 · 12 citations
- HintedBT: Augmenting Back-Translation with Quality and Transliteration HintsSahana Ramnath, Melvin Johnson, Abhirut Gupta, Aravindan RaghuveerEMNLP 2021 · 7 citations
- Enhancing Cross-Lingual Transfer through Reversible Transliteration: A Huffman-Based Approach for Low-Resource LanguagesWenhao Zhuang, Yuan Sun, Xiaobing ZhaoACL 2025 · 1 citation
- On the Generalization of Handwritten Text Recognition ModelsCarlos Garrido-Munoz, Jorge Calvo-ZaragozaCVPR 2025
