Filling the Gaps in Ancient Akkadian Texts: A Masked Language Modelling Approach
Koren Lazar, Benny Saret, Asaf Yehudai, Wayne Horowitz, Nathan Wasserman, Gabriel Stanovsky
摘要
We present models which complete missing text given transliterations of ancient Mesopotamian documents, originally written on cuneiform clay tablets (2500 BCE -100 CE). Due to the tablets' deterioration, scholars often rely on contextual cues to manually fill in missing parts in the text in a subjective and time-consuming process. We identify that this challenge can be formulated as a masked language modelling task, used mostly as a pretraining objective for contextualized language models. Following, we develop several architectures focusing on the Akkadian language, the lingua franca of the time. We find that despite data scarcity (1M tokens) we can achieve state of the art performance on missing tokens prediction (89% hit@5) using a greedy decoding scheme and pretraining on data from other languages and different time periods. Finally, we conduct human evaluations showing the applicability of our models in assisting experts to transcribe texts in extinct languages.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Reviving Cultural Heritage: A Novel Approach for Comprehensive Historical Document RestorationYuyi Zhang, Peirong Zhang, Zhenhua Yang, Pengyu Yan 等ACL 2025 · 被引用 5 次
- PhiloGPT: A Philology-Oriented Large Language Model for Ancient Chinese Manuscripts with Dunhuang as Case StudyYuqing Zhang, Baoyi He, Yihan Chen, Hangqi Li 等EMNLP 2024 · 被引用 1 次
- Draft, Verify, Restore: Self-Refining Historical Inscription Restoration with a Unified MLLMYuyi Zhang, Junle Liu, Peirong Zhang, Jianliang Liu 等ACL 2026
它引用的顶会 Paper1
相关 Paper
- UniLMv2: Pseudo-Masked Language Models for Unified Language Model Pre-TrainingHangbo Bao, Li Dong, Furu Wei, Wenhui Wang 等ICML 2020 · 被引用 423 次
- Blank Language ModelsTianxiao Shen, Victor Quach, Regina Barzilay, Tommi S. JaakkolaEMNLP 2020 · 被引用 8 次
- Learned Meta-Tokens for Language ModelingAlok N. Shah, Khush Gupta, Keshav Ramji, Pratik ChaudhariICLR 2026 · 被引用 2 次
- Script, Language, and Labels: Overcoming Three Discrepancies for Low-Resource Language SpecializationJaeseong Lee, Dohyeon Lee, Seung-won HwangAAAI 2023 · 被引用 1 次
- ExLM: Rethinking the Impact of [MASK] Tokens in Masked Language ModelsKangjie Zheng, Junwei Yang, Siyue Liang, Bin Feng 等ICML 2025
