Quantifying Character Similarity with Vision Transformers
Xinmei Yang, Abhishek Arora, Shao-Yu Jheng, Melissa Dell
摘要
Record linkage is a bedrock of quantitative social science, as analyses often require linking data from multiple, noisy sources. Off-the-shelf string matching methods are widely used, as they are straightforward and cheap to implement and scale. Not all character substitutions are equally probable, and for some settings there are widely used handcrafted lists denoting which string substitutions are more likely, that improve the accuracy of string matching. However, such lists do not exist for many settings, skewing research with linked datasets towards a few high-resource contexts that are not representative of the diversity of human societies. This study develops an extensible way to measure character substitution costs for OCR’ed documents, by employing large-scale self-supervised training of vision transformers (ViT) with augmented digital fonts. For each language written with the CJK script, we contrastively learn a metric space where different augmentations of the same character are represented nearby. In this space, homoglyphic characters - those with similar appearance such as “O” and “0” - have similar vector representations. Using the cosine distance between characters’ representations as the substitution cost in an edit distance matching algorithm significantly improves record linkage compared to other widely used string matching methods, as OCR errors tend to be homoglyphic in nature. Homoglyphs can plausibly capture character visual similarity across any script, including low-resource settings. We illustrate this by creating homoglyph sets for 3,000 year old ancient Chinese characters, which are highly pictorial. Fascinatingly, a ViT is able to capture relationships in how different abstract concepts were conceptualized by ancient societies, that have been noted in the archaeological literature.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna 等NeurIPS 2020 · 被引用 7,049 次
- An Empirical Study of Training Self-Supervised Vision TransformersXinlei Chen, Saining Xie, Kaiming HeICCV 2021 · 被引用 2,340 次
- Efficient OCR for Building a Diverse Digital HistoryJacob Carlson, Tom Bryan, Melissa DellACL 2024 · 被引用 6 次
相关 Paper
- Beyond Scalars: Concept-Based Alignment Analysis in Vision TransformersJohanna Vielhaben, Dilyara Bareeva, Jim Berend, Wojciech Samek 等NeurIPS 2025 · 被引用 11 次
- UPOCR: Towards Unified Pixel-Level OCR InterfaceDezhi Peng, Zhenhua Yang, Jiaxin Zhang, Chongyu Liu 等ICML 2024 · 被引用 14 次
- Enhancing Multimodal Large Language Models for Ancient Chinese Character Evolution Analysis via Glyph-Driven Fine-TuningRui Song, Lida Shi, Ruihua Qi, Yingji Li 等ACL 2026 · 被引用 1 次
- XMP-Font: Self-Supervised Cross-Modality Pre-training for Few-Shot Font GenerationWei Liu, Fangyue Liu, Fei Ding, Qian He 等CVPR 2022 · 被引用 64 次
- Specializing Large Models for Oracle Bone Script Interpretation via Component-Grounded Multimodal Knowledge AugmentationJianing Zhang, Runan Li, Honglin Pang, Ding Xia 等ACL 2026 · 被引用 1 次
