Massively Multilingual Joint Segmentation and Glossing
Michael Ginn, Lindia Tjuatja, Enora Rice, Ali Marashian, Maria R. Valentini, Jasmine Xu, Graham Neubig, Alexis Palmer
摘要
Automated interlinear gloss prediction with neural networks is a promising approach to accelerate language documentation efforts. However, while state-of-the-art models like GLOSSLM (Ginn et al., 2024b) achieve high scores on glossing benchmarks, user studies with linguists have found critical barriers to the usefulness of such models in real-world scenarios (Rice et al., 2025) . In particular, existing models typically generate morpheme-level glosses but assign them to whole words without predicting the actual morpheme boundaries, making the predictions less interpretable and thus untrustworthy to human annotators. We conduct the first study on neural models that jointly predict interlinear glosses and the corresponding morphological segmentation from raw text. We run experiments to determine the optimal way to train models that balance segmentation and glossing accuracy, as well as the alignment between the two tasks. We extend the training corpus of GLOSSLM and pretrain POLYGLOSS, a family of seq2seq multilingual models for joint segmentation and glossing that outperforms GLOSSLM on glossing and beats various open-source LLMs on segmentation, glossing, and alignment. In addition, we demonstrate that POLYGLOSS can be quickly adapted to a new dataset via low-rank adaptation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- IGT2P: From Interlinear Glossed Texts to ParadigmsSarah R. Moeller, Ling Liu, Changbing Yang, Katharina Kann 等EMNLP 2020 · 被引用 13 次
- GrammaMT: Improving Machine Translation with Grammar-Informed In-Context LearningRita Ramos, Everlyn Asiko Chimoto, Maartje ter Hoeve, Natalie SchluterACL 2025 · 被引用 10 次
- Wav2Gloss: Generating Interlinear Glossed Text from SpeechTaiqi He, Kwanghee Choi, Lindia Tjuatja, Nathaniel R. Robinson 等ACL 2024 · 被引用 1 次
- GlossLM: A Massively Multilingual Corpus and Pretrained Model for Interlinear Glossed TextMichael Ginn, Lindia Tjuatja, Taiqi He, Enora Rice 等EMNLP 2024 · 被引用 1 次
相关 Paper
- Interdisciplinary Research in Conversation: A Case Study in Computational Morphology for Language DocumentationEnora Rice, Katharina von der Wense, Alexis PalmerEMNLP 2025
- TAMS: Translation-Assisted Morphological SegmentationEnora Rice, Ali Marashian, Luke Gessler, Alexis Palmer 等ACL 2024
- To POS Tag or Not to POS Tag: The Impact of POS Tags on Morphological Learning in Low-Resource SettingsSarah R. Moeller, Ling Liu, Mans HuldenACL 2021
- TASTE: Text-Aligned Speech Tokenization and Embedding for Spoken Language ModelingLiang-Hsuan Tseng, Yi-Chang Chen, Kuan Yi Lee, Da-shan Shiu 等ICLR 2026 · 被引用 26 次
- Instruction-guided Multi-Granularity Segmentation and Captioning with Large Multimodal ModelXu Yuan, Li Zhou, Zenghui Sun, Zikun Zhou 等AAAI 2025 · 被引用 1 次
