Sequence Models for Document Structure Identification in an Undeciphered Script
Logan Born, M. Willis Monroe, Kathryn Kelley, Anoop Sarkar
摘要
This work describes the first thorough analysis of “header” signs in proto-Elamite, an undeciphered script from 3100-2900 BCE. Headers are a category of signs which have been provisionally identified through painstaking manual analysis of this script by domain experts. We use unsupervised neural and statistical sequence modeling techniques to provide new and independent evidence for the existence of headers, without supervision from domain experts. Having affirmed the existence of headers as a legitimate structural feature, we next arrive at a richer understanding of their possible meaning and purpose by (i) examining which features predict their presence; (ii) identifying correlations between these features and other document properties; and (iii) examining cases where these features predict the presence of a header in texts where domain experts do not expect one (or vice versa). We provide more concrete processes for labeling headers in this corpus and a clearer justification for existing intuitions about document structure in proto-Elamite.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
相关 Paper
- ProtoSnap: Prototype Alignment For Cuneiform SignsRachel Mikulinsky, Morris Alper, Shai Gordin, Enrique Jiménez 等ICLR 2025
- Unsupervised Boundary-Aware Language Model Pretraining for Chinese Sequence LabelingPeijie Jiang, Dingkun Long, Yanzhao Zhang, Pengjun Xie 等EMNLP 2022 · 被引用 9 次
- Chapter Captor: Text Segmentation in NovelsCharuta Pethe, Allen Kim, Steven SkienaEMNLP 2020 · 被引用 18 次
- Filling the Gaps in Ancient Akkadian Texts: A Masked Language Modelling ApproachKoren Lazar, Benny Saret, Asaf Yehudai, Wayne Horowitz 等EMNLP 2021 · 被引用 15 次
- Asking without Telling: Exploring Latent Ontologies in Contextual RepresentationsJulian Michael, Jan A. Botha, Ian TenneyEMNLP 2020 · 被引用 3 次
