Multiple Sources are Better Than One: Incorporating External Knowledge in Low-Resource Glossing
Changbing Yang, Garrett Nicolai, Miikka Silfverberg
Abstract
In this paper, we address the data scarcity problem in automatic data-driven glossing for low-resource languages by coordinating multiple sources of linguistic expertise. We enhance models by incorporating both tokenlevel and sentence-level translations, utilizing the extensive linguistic capabilities of modern LLMs, and incorporating available dictionary resources. Our enhancements lead to an average absolute improvement of 5%-points in word-level accuracy over the previous state of the art on a typologically diverse dataset spanning six low-resource languages. The improvements are particularly noticeable for the lowest-resourced language Gitksan, where we achieve a 10%-point improvement. Furthermore, in a simulated ultra-low resource setting for the same six languages, training on fewer than 100 glossed sentences, we establish an average 10%-point improvement in word-level accuracy over the previous state-of-the-art system.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0e5374ac-e7b5-4e03-9a4e-4b5b12a08891Cited by top-tier papers3
- Massively Multilingual Joint Segmentation and GlossingMichael Ginn, Lindia Tjuatja, Enora Rice, Ali Marashian et al.ACL 2026 · 2 citations
- Interdisciplinary Research in Conversation: A Case Study in Computational Morphology for Language DocumentationEnora Rice, Katharina von der Wense, Alexis PalmerEMNLP 2025
- LingGym: How Far Are LLMs from Thinking Like Field Linguists?Changbing Yang, Franklin Ma, Freda Shi, Jian ZhuEMNLP 2025
Builds on2
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- Improving Low-Resource Languages in Pre-Trained Multilingual Language ModelsViktor Hangya, Hossain Shaikh Saadi, Alexander FraserEMNLP 2022 · 17 citations
Related papers
- Is linguistically-motivated data augmentation worth it?Ray Groshan, Michael Ginn, Alexis PalmerACL 2025
- Neural Machine Translation Methods for Translating Text to Sign Language GlossesDele Zhu, Vera Czehmann, Eleftherios AvramidisACL 2023 · 12 citations
- GrammaMT: Improving Machine Translation with Grammar-Informed In-Context LearningRita Ramos, Everlyn Asiko Chimoto, Maartje ter Hoeve, Natalie SchluterACL 2025 · 10 citations
- Small Data, Big Impact: Leveraging Minimal Data for Effective Machine TranslationJean Maillard, Cynthia Gao, Elahe Kalbassi, Kaushik Ram Sadagopan et al.ACL 2023 · 3 citations
- Can LLMs Really Learn to Translate a Low-Resource Language from One Grammar Book?Seth Aycock, David Stap, Di Wu, Christof Monz et al.ICLR 2025
