GlossLM: A Massively Multilingual Corpus and Pretrained Model for Interlinear Glossed Text
Michael Ginn, Lindia Tjuatja, Taiqi He, Enora Rice, Graham Neubig, Alexis Palmer, Lori S. Levin
摘要
Language documentation projects often involve the creation of annotated text in a format such as interlinear glossed text (IGT), which captures fine-grained morphosyntactic analyses in a morpheme-by-morpheme format. However, there are few existing resources providing large amounts of standardized, easily accessible IGT data, limiting their applicability to linguistic research, and making it difficult to use such data in NLP modeling. We compile the largest existing corpus of IGT data from a variety of sources, covering over 450k examples across 1.8k languages, to enable research on crosslingual transfer and IGT generation. We normalize much of our data to follow a standard set of labels across languages. Furthermore, we explore the task of automatically generating IGT in order to aid documentation projects. As many languages lack sufficient monolingual data, we pretrain a large multilingual model on our corpus. We demonstrate the utility of this model by finetuning it on monolingual corpora, outperforming SOTA models by up to 6.6%. Our pretrained model and dataset are available on Hugging Face. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Read it in Two Steps: Translating Extremely Low-Resource Languages with Code-Augmented Grammar BooksChen Zhang, Jiuheng Lin, Xiao Liu, Zekai Zhang 等ACL 2025 · 被引用 5 次
- Massively Multilingual Joint Segmentation and GlossingMichael Ginn, Lindia Tjuatja, Enora Rice, Ali Marashian 等ACL 2026 · 被引用 2 次
- Interdisciplinary Research in Conversation: A Case Study in Computational Morphology for Language DocumentationEnora Rice, Katharina von der Wense, Alexis PalmerEMNLP 2025
它引用的顶会 Paper4
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- MasakhaNER 2.0: Africa-centric Transfer Learning for Named Entity RecognitionDavid Ifeoluwa Adelani, Graham Neubig, Sebastian Ruder, Shruti Rijhwani 等EMNLP 2022 · 被引用 46 次
- IGT2P: From Interlinear Glossed Texts to ParadigmsSarah R. Moeller, Ling Liu, Changbing Yang, Katharina Kann 等EMNLP 2020 · 被引用 13 次
- When Is Multilinguality a Curse? Language Modeling for 250 High- and Low-Resource LanguagesTyler A. Chang, Catherine Arnett, Zhuowen Tu, Ben BergenEMNLP 2024 · 被引用 12 次
相关 Paper
- Wav2Gloss: Generating Interlinear Glossed Text from SpeechTaiqi He, Kwanghee Choi, Lindia Tjuatja, Nathaniel R. Robinson 等ACL 2024 · 被引用 1 次
- An Expanded Massive Multilingual Dataset for High-Performance Language Technologies (HPLT)Laurie Burchell, Ona De Gibert Bonet, Nikolay Arefyev, Mikko Aulamo 等ACL 2025
- Glot500: Scaling Multilingual Corpora and Language Models to 500 LanguagesAyyoob Imani, Peiqin Lin, Amir Hossein Kargaran, Silvia Severini 等ACL 2023 · 被引用 14 次
- XGLUE: A New Benchmark Datasetfor Cross-lingual Pre-training, Understanding and GenerationYaobo Liang, Nan Duan, Yeyun Gong, Ning Wu 等EMNLP 2020 · 被引用 232 次
- Common Corpus: The Largest Collection of Ethical Data for LLM Pre-TrainingPierre-Carl Langlais, Pavel Chizhov, Catherine Arnett, Carlos Rosas Hinostroza 等ICLR 2026 · 被引用 22 次
