LogogramNLP: Comparing Visual and Textual Representations of Ancient Logographic Writing Systems for NLP
Danlu Chen, Freda Shi, Aditi Agarwal, Jacobo Myerston, Taylor Berg-Kirkpatrick
Abstract
Standard natural language processing (NLP) pipelines operate on symbolic representations of language, which typically consist of sequences of discrete tokens. However, creating an analogous representation for ancient logographic writing systems is an extremely laborintensive process that requires expert knowledge. At present, a large portion of logographic data persists in a purely visual form due to the absence of transcription-this issue poses a bottleneck for researchers seeking to apply NLP toolkits to study ancient logographic languages: most of the relevant data are images of writing. This paper investigates whether direct processing of visual representations of language offers a potential solution. We introduce LogogramNLP, the first benchmark enabling NLP analysis of ancient logographic languages, featuring both transcribed and visual datasets for four writing systems along with annotations for tasks like classification, translation, and parsing. Our experiments compare systems that employ recent visual and text encoding strategies as backbones. The results demonstrate that visual representations outperform textual representations for some investigated tasks, suggesting that visual processing pipelines may unlock a large amount of cultural heritage data of logographic languages for NLP-based analyses. Data and code are available at https: //logogramNLP.github.io/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f048afd5-af83-4200-aed5-79515fc2489dBuilds on6
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- Language Modelling with PixelsPhillip Rust, Jonas F. Lotz, Emanuele Bugliarello, Elizabeth Salesky et al.ICLR 2023 · 17 citations
- Multilingual Pixel Representations for Translation and Effective Cross-lingual TransferElizabeth Salesky, Neha Verma, Philipp Koehn, Matt PostEMNLP 2023
- CLIPPO: Image-and-Language Understanding from Pixels OnlyMichael Tschannen, Basil Mustafa, Neil HoulsbyCVPR 2023
- How Good is Your Tokenizer? On the Monolingual Performance of Multilingual Language ModelsPhillip Rust, Jonas Pfeiffer, Ivan Vulic, Sebastian Ruder et al.ACL 2021
Related papers
- NusaAksara: A Multimodal and Multilingual Benchmark for Preserving Indonesian Indigenous ScriptsMuhammad Farid Adilazuarda, Musa Izzanardi Wijanarko, Lucky Susanto, Khumaisa Nur'aini et al.ACL 2025
- BoYaEval: Evaluating Multimodal Large Language Models on Understanding Ancient Chinese Musical ScoresJiajia Li, Weizhi Xue, Yao Yao, Qiwei Li et al.ACL 2026
- Coding Textual Inputs Boosts the Accuracy of Neural NetworksAbdul Rafae Khan, Jia Xu, Weiwei SunEMNLP 2020 · 2 citations
- AncientBench: Towards Comprehensive Evaluation on Excavated and Transmitted Chinese CorporaZhihan Zhou, Daqian Shi, Rui Song, Lida Shi et al.AAAI 2026 · 1 citation
- PhiloGPT: A Philology-Oriented Large Language Model for Ancient Chinese Manuscripts with Dunhuang as Case StudyYuqing Zhang, Baoyi He, Yihan Chen, Hangqi Li et al.EMNLP 2024 · 1 citation
