Filling the Gaps in Ancient Akkadian Texts: A Masked Language Modelling Approach
Koren Lazar, Benny Saret, Asaf Yehudai, Wayne Horowitz, Nathan Wasserman, Gabriel Stanovsky
Abstract
We present models which complete missing text given transliterations of ancient Mesopotamian documents, originally written on cuneiform clay tablets (2500 BCE -100 CE). Due to the tablets' deterioration, scholars often rely on contextual cues to manually fill in missing parts in the text in a subjective and time-consuming process. We identify that this challenge can be formulated as a masked language modelling task, used mostly as a pretraining objective for contextualized language models. Following, we develop several architectures focusing on the Akkadian language, the lingua franca of the time. We find that despite data scarcity (1M tokens) we can achieve state of the art performance on missing tokens prediction (89% hit@5) using a greedy decoding scheme and pretraining on data from other languages and different time periods. Finally, we conduct human evaluations showing the applicability of our models in assisting experts to transcribe texts in extinct languages.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Reviving Cultural Heritage: A Novel Approach for Comprehensive Historical Document RestorationYuyi Zhang, Peirong Zhang, Zhenhua Yang, Pengyu Yan et al.ACL 2025 · 5 citations
- PhiloGPT: A Philology-Oriented Large Language Model for Ancient Chinese Manuscripts with Dunhuang as Case StudyYuqing Zhang, Baoyi He, Yihan Chen, Hangqi Li et al.EMNLP 2024 · 1 citation
- Draft, Verify, Restore: Self-Refining Historical Inscription Restoration with a Unified MLLMYuyi Zhang, Junle Liu, Peirong Zhang, Jianliang Liu et al.ACL 2026
Builds on1
Related papers
- UniLMv2: Pseudo-Masked Language Models for Unified Language Model Pre-TrainingHangbo Bao, Li Dong, Furu Wei, Wenhui Wang et al.ICML 2020 · 423 citations
- Blank Language ModelsTianxiao Shen, Victor Quach, Regina Barzilay, Tommi S. JaakkolaEMNLP 2020 · 8 citations
- Learned Meta-Tokens for Language ModelingAlok N. Shah, Khush Gupta, Keshav Ramji, Pratik ChaudhariICLR 2026 · 2 citations
- Script, Language, and Labels: Overcoming Three Discrepancies for Low-Resource Language SpecializationJaeseong Lee, Dohyeon Lee, Seung-won HwangAAAI 2023 · 1 citation
- ExLM: Rethinking the Impact of [MASK] Tokens in Masked Language ModelsKangjie Zheng, Junwei Yang, Siyue Liang, Bin Feng et al.ICML 2025
