Massively Multilingual Joint Segmentation and Glossing
Michael Ginn, Lindia Tjuatja, Enora Rice, Ali Marashian, Maria R. Valentini, Jasmine Xu, Graham Neubig, Alexis Palmer
Abstract
Automated interlinear gloss prediction with neural networks is a promising approach to accelerate language documentation efforts. However, while state-of-the-art models like GLOSSLM (Ginn et al., 2024b) achieve high scores on glossing benchmarks, user studies with linguists have found critical barriers to the usefulness of such models in real-world scenarios (Rice et al., 2025) . In particular, existing models typically generate morpheme-level glosses but assign them to whole words without predicting the actual morpheme boundaries, making the predictions less interpretable and thus untrustworthy to human annotators. We conduct the first study on neural models that jointly predict interlinear glosses and the corresponding morphological segmentation from raw text. We run experiments to determine the optimal way to train models that balance segmentation and glossing accuracy, as well as the alignment between the two tasks. We extend the training corpus of GLOSSLM and pretrain POLYGLOSS, a family of seq2seq multilingual models for joint segmentation and glossing that outperforms GLOSSLM on glossing and beats various open-source LLMs on segmentation, glossing, and alignment. In addition, we demonstrate that POLYGLOSS can be quickly adapted to a new dataset via low-rank adaptation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 59894cc2-59e3-414e-802d-48fcd7ec0f32Builds on10
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- IGT2P: From Interlinear Glossed Texts to ParadigmsSarah R. Moeller, Ling Liu, Changbing Yang, Katharina Kann et al.EMNLP 2020 · 13 citations
- GrammaMT: Improving Machine Translation with Grammar-Informed In-Context LearningRita Ramos, Everlyn Asiko Chimoto, Maartje ter Hoeve, Natalie SchluterACL 2025 · 10 citations
- Wav2Gloss: Generating Interlinear Glossed Text from SpeechTaiqi He, Kwanghee Choi, Lindia Tjuatja, Nathaniel R. Robinson et al.ACL 2024 · 1 citation
- GlossLM: A Massively Multilingual Corpus and Pretrained Model for Interlinear Glossed TextMichael Ginn, Lindia Tjuatja, Taiqi He, Enora Rice et al.EMNLP 2024 · 1 citation
Related papers
- Interdisciplinary Research in Conversation: A Case Study in Computational Morphology for Language DocumentationEnora Rice, Katharina von der Wense, Alexis PalmerEMNLP 2025
- TAMS: Translation-Assisted Morphological SegmentationEnora Rice, Ali Marashian, Luke Gessler, Alexis Palmer et al.ACL 2024
- To POS Tag or Not to POS Tag: The Impact of POS Tags on Morphological Learning in Low-Resource SettingsSarah R. Moeller, Ling Liu, Mans HuldenACL 2021
- TASTE: Text-Aligned Speech Tokenization and Embedding for Spoken Language ModelingLiang-Hsuan Tseng, Yi-Chang Chen, Kuan Yi Lee, Da-shan Shiu et al.ICLR 2026 · 26 citations
- Instruction-guided Multi-Granularity Segmentation and Captioning with Large Multimodal ModelXu Yuan, Li Zhou, Zenghui Sun, Zikun Zhou et al.AAAI 2025 · 1 citation
