Efficient OCR for Building a Diverse Digital History
Jacob Carlson, Tom Bryan, Melissa Dell
Abstract
Many users consult digital archives daily, but the information they can access is unrepresentative of the diversity of documentary history. The sequence-to-sequence architecture typically used for optical character recognition (OCR) -which jointly learns a vision and language model -is poorly extensible to lowresource document collections, as learning a language-vision model requires extensive labeled sequences and compute. This study models OCR as a character level image retrieval problem, using a contrastively trained vision encoder. Because the model only learns characters' visual features, it is more sample efficient and extensible than existing architectures, enabling accurate OCR in settings where existing solutions fail. Crucially, it opens new avenues for community engagement in making digital history more representative of documentary history.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Quantifying Character Similarity with Vision TransformersXinmei Yang, Abhishek Arora, Shao-Yu Jheng, Melissa DellEMNLP 2023 · 2 citations
- CalligraphicOCR for Chinese Calligraphy RecognitionXiaoyi Bao, Zhongqing Wang, Jinghang Gu, Chu-Ren HuangEMNLP 2025 · 1 citation
Builds on13
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- An Empirical Study of Training Self-Supervised Vision TransformersXinlei Chen, Saining Xie, Kaiming HeICCV 2021 · 2,340 citations
Related papers
- Rethinking Genomic Modeling Through Optical Character RecognitionHongxin Xiang, Pengsen Ma, Yunkang Cao, Di Yu et al.ICML 2026
- Sequence-to-Sequence Contrastive Learning for Text RecognitionAviad Aberdam, Ron Litman, Shahar Tsiper, Oron Anschel et al.CVPR 2021
- From Token to Word: OCR Token Evolution via Contrastive Learning and Semantic Matching for Text-VQAZan-Xia Jin, Mike Zheng Shou, Fang Zhou, Satoshi Tsutsui et al.ACM MM 2022 · 11 citations
- ModernVBERT: Towards Smaller Visual Document RetrieversPaul Teiletche, Quentin Macé, Max Conti, António Loison et al.ICML 2026 · 17 citations
- Post-OCR Document Correction with Large Ensembles of Character Sequence-to-Sequence ModelsJuan Antonio Ramirez-Orta, Eduardo Xamena, Ana Gabriela Maguitman, Evangelos E. Milios et al.AAAI 2022 · 20 citations
