LogogramNLP: Comparing Visual and Textual Representations of Ancient Logographic Writing Systems for NLP
Danlu Chen, Freda Shi, Aditi Agarwal, Jacobo Myerston, Taylor Berg-Kirkpatrick
摘要
Standard natural language processing (NLP) pipelines operate on symbolic representations of language, which typically consist of sequences of discrete tokens. However, creating an analogous representation for ancient logographic writing systems is an extremely laborintensive process that requires expert knowledge. At present, a large portion of logographic data persists in a purely visual form due to the absence of transcription-this issue poses a bottleneck for researchers seeking to apply NLP toolkits to study ancient logographic languages: most of the relevant data are images of writing. This paper investigates whether direct processing of visual representations of language offers a potential solution. We introduce LogogramNLP, the first benchmark enabling NLP analysis of ancient logographic languages, featuring both transcribed and visual datasets for four writing systems along with annotations for tasks like classification, translation, and parsing. Our experiments compare systems that employ recent visual and text encoding strategies as backbones. The results demonstrate that visual representations outperform textual representations for some investigated tasks, suggesting that visual processing pipelines may unlock a large amount of cultural heritage data of logographic languages for NLP-based analyses. Data and code are available at https: //logogramNLP.github.io/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- Language Modelling with PixelsPhillip Rust, Jonas F. Lotz, Emanuele Bugliarello, Elizabeth Salesky 等ICLR 2023 · 被引用 17 次
- Multilingual Pixel Representations for Translation and Effective Cross-lingual TransferElizabeth Salesky, Neha Verma, Philipp Koehn, Matt PostEMNLP 2023
- CLIPPO: Image-and-Language Understanding from Pixels OnlyMichael Tschannen, Basil Mustafa, Neil HoulsbyCVPR 2023
- How Good is Your Tokenizer? On the Monolingual Performance of Multilingual Language ModelsPhillip Rust, Jonas Pfeiffer, Ivan Vulic, Sebastian Ruder 等ACL 2021
相关 Paper
- NusaAksara: A Multimodal and Multilingual Benchmark for Preserving Indonesian Indigenous ScriptsMuhammad Farid Adilazuarda, Musa Izzanardi Wijanarko, Lucky Susanto, Khumaisa Nur'aini 等ACL 2025
- BoYaEval: Evaluating Multimodal Large Language Models on Understanding Ancient Chinese Musical ScoresJiajia Li, Weizhi Xue, Yao Yao, Qiwei Li 等ACL 2026
- Coding Textual Inputs Boosts the Accuracy of Neural NetworksAbdul Rafae Khan, Jia Xu, Weiwei SunEMNLP 2020 · 被引用 2 次
- AncientBench: Towards Comprehensive Evaluation on Excavated and Transmitted Chinese CorporaZhihan Zhou, Daqian Shi, Rui Song, Lida Shi 等AAAI 2026 · 被引用 1 次
- PhiloGPT: A Philology-Oriented Large Language Model for Ancient Chinese Manuscripts with Dunhuang as Case StudyYuqing Zhang, Baoyi He, Yihan Chen, Hangqi Li 等EMNLP 2024 · 被引用 1 次
