Improving Low-Resource Morphological Inflection via Self-Supervised Objectives
Adam Wiemerslage, Katharina von der Wense
Abstract
Self-supervised objectives have driven major advances in NLP by leveraging large-scale unlabeled data, but such resources are scarce for many of the world's languages. Surprisingly, they have not been explored much for character-level tasks, where smaller amounts of data have the potential to be beneficial. We investigate the effectiveness of self-supervised auxiliary tasks for morphological inflection -- a character-level task highly relevant for language documentation -- in extremely low-resource settings, training encoder-decoder transformers for 19 languages and 13 auxiliary objectives. Autoencoding yields the best performance when unlabeled data is very limited, while character masked language modeling (CMLM) becomes more effective as data availability increases. Though objectives with stronger inductive biases influence model predictions intuitively, they rarely outperform standard CMLM. However, sampling masks based on known morpheme boundaries consistently improves performance, highlighting a promising direction for low-resource morphological modeling.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e569a1ea-46bd-494f-b5d3-c27cf1b613b0Builds on4
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- Cross-Task Generalization via Natural Language Crowdsourcing InstructionsSwaroop Mishra, Daniel Khashabi, Chitta Baral, Hannaneh HajishirziACL 2022 · 887 citations
- IGT2P: From Interlinear Glossed Texts to ParadigmsSarah R. Moeller, Ling Liu, Changbing Yang, Katharina Kann et al.EMNLP 2020 · 13 citations
- Getting The Most Out of Your Training Data: Exploring Unsupervised Tasks for Morphological InflectionAbhishek Purushothama, Adam Wiemerslage, Katharina von der WenseEMNLP 2024
Related papers
- TAMS: Translation-Assisted Morphological SegmentationEnora Rice, Ali Marashian, Luke Gessler, Alexis Palmer et al.ACL 2024
- MYTE: Morphology-Driven Byte Encoding for Better and Fairer Multilingual Language ModelingTomasz Limisiewicz, Terra Blevins, Hila Gonen, Orevaoghene Ahia et al.ACL 2024 · 1 citation
- Learning to Learn Morphological Inflection for Resource-Poor LanguagesKatharina Kann, Samuel R. Bowman, Kyunghyun ChoAAAI 2020 · 9 citations
- Hints on the data for language modeling of synthetic languages with transformersRodolfo Zevallos, Núria BelACL 2023 · 2 citations
- Weakly Supervised POS Taggers Perform Poorly on Truly Low-Resource LanguagesKatharina Kann, Ophélie Lacroix, Anders SøgaardAAAI 2020 · 21 citations
