ProtoSnap: Prototype Alignment For Cuneiform Signs
Rachel Mikulinsky, Morris Alper, Shai Gordin, Enrique Jiménez, Yoram Cohen, Hadar Averbuch-Elor
Abstract
The cuneiform writing system served as the medium for transmitting knowledge in the ancient Near East for a period of over three thousand years. Cuneiform signs have a complex internal structure which is the subject of expert paleographic analysis, as variations in sign shapes bear witness to historical developments and transmission of writing and culture over time. However, prior automated techniques mostly treat sign types as categorical and do not explicitly model their highly varied internal configurations. In this work, we present an unsupervised approach for recovering the fine-grained internal configuration of cuneiform signs by leveraging powerful generative models and the appearance and structure of prototype font images as priors. Our approach, Pro-toSnap, enforces structural consistency on matches found with deep image features to estimate the diverse configurations of cuneiform characters, snapping a skeleton-based template to photographed cuneiform signs. We provide a new benchmark of expert annotations and evaluate our method on this task. Our evaluation shows that our approach succeeds in aligning prototype skeletons to a wide variety of cuneiform signs. Moreover, we show that conditioning on structures produced by our method allows for generating synthetic data with correct structural configurations, significantly boosting the performance of cuneiform sign recognition beyond existing techniques, in particular over rare signs. Our code, data, and trained models are available at the project page: https://tau-vailab.github.io/ProtoSnap/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on5
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- SwinTextSpotter: Scene Text Spotting via Better Synergy between Text Detection and Text RecognitionMingxin Huang, Yuliang Liu, Zhenghao Peng, Chongyu Liu et al.CVPR 2022 · 150 citations
- Recurrent Temporal Revision Graph NetworksYizhou Chen, Anxiang Zeng, Qingtao Yu, Kerui Zhang et al.NeurIPS 2023 · 6 citations
- From Synthetic to Real: Unsupervised Domain Adaptation for Animal Pose EstimationChen Li, Gim Hee LeeCVPR 2021
Related papers
- AGTGAN: Unpaired Image Translation for Photographic Ancient Character GenerationHongxiang Huang, Daihui Yang, Gang Dai, Zhen Han et al.ACM MM 2022 · 31 citations
- Sequence Models for Document Structure Identification in an Undeciphered ScriptLogan Born, M. Willis Monroe, Kathryn Kelley, Anoop SarkarEMNLP 2022 · 4 citations
- Filling the Gaps in Ancient Akkadian Texts: A Masked Language Modelling ApproachKoren Lazar, Benny Saret, Asaf Yehudai, Wayne Horowitz et al.EMNLP 2021 · 15 citations
- V-Oracle: Making Progressive Reasoning in Deciphering Oracle Bones for You and MeRunqi Qiao, Qiuna Tan, Guanting Dong, Minhui Wu et al.ACL 2025 · 8 citations
- Scalable Font Reconstruction with Dual Latent ManifoldsNikita Srivatsan, Si Wu, Jonathan T. Barron, Taylor Berg-KirkpatrickEMNLP 2021 · 1 citation
