ProtoSnap: Prototype Alignment For Cuneiform Signs
Rachel Mikulinsky, Morris Alper, Shai Gordin, Enrique Jiménez, Yoram Cohen, Hadar Averbuch-Elor
摘要
The cuneiform writing system served as the medium for transmitting knowledge in the ancient Near East for a period of over three thousand years. Cuneiform signs have a complex internal structure which is the subject of expert paleographic analysis, as variations in sign shapes bear witness to historical developments and transmission of writing and culture over time. However, prior automated techniques mostly treat sign types as categorical and do not explicitly model their highly varied internal configurations. In this work, we present an unsupervised approach for recovering the fine-grained internal configuration of cuneiform signs by leveraging powerful generative models and the appearance and structure of prototype font images as priors. Our approach, Pro-toSnap, enforces structural consistency on matches found with deep image features to estimate the diverse configurations of cuneiform characters, snapping a skeleton-based template to photographed cuneiform signs. We provide a new benchmark of expert annotations and evaluate our method on this task. Our evaluation shows that our approach succeeds in aligning prototype skeletons to a wide variety of cuneiform signs. Moreover, we show that conditioning on structures produced by our method allows for generating synthetic data with correct structural configurations, significantly boosting the performance of cuneiform sign recognition beyond existing techniques, in particular over rare signs. Our code, data, and trained models are available at the project page: https://tau-vailab.github.io/ProtoSnap/
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
- SwinTextSpotter: Scene Text Spotting via Better Synergy between Text Detection and Text RecognitionMingxin Huang, Yuliang Liu, Zhenghao Peng, Chongyu Liu 等CVPR 2022 · 被引用 150 次
- Recurrent Temporal Revision Graph NetworksYizhou Chen, Anxiang Zeng, Qingtao Yu, Kerui Zhang 等NeurIPS 2023 · 被引用 6 次
- From Synthetic to Real: Unsupervised Domain Adaptation for Animal Pose EstimationChen Li, Gim Hee LeeCVPR 2021
相关 Paper
- AGTGAN: Unpaired Image Translation for Photographic Ancient Character GenerationHongxiang Huang, Daihui Yang, Gang Dai, Zhen Han 等ACM MM 2022 · 被引用 31 次
- Sequence Models for Document Structure Identification in an Undeciphered ScriptLogan Born, M. Willis Monroe, Kathryn Kelley, Anoop SarkarEMNLP 2022 · 被引用 4 次
- Filling the Gaps in Ancient Akkadian Texts: A Masked Language Modelling ApproachKoren Lazar, Benny Saret, Asaf Yehudai, Wayne Horowitz 等EMNLP 2021 · 被引用 15 次
- V-Oracle: Making Progressive Reasoning in Deciphering Oracle Bones for You and MeRunqi Qiao, Qiuna Tan, Guanting Dong, Minhui Wu 等ACL 2025 · 被引用 8 次
- Scalable Font Reconstruction with Dual Latent ManifoldsNikita Srivatsan, Si Wu, Jonathan T. Barron, Taylor Berg-KirkpatrickEMNLP 2021 · 被引用 1 次
