Bootstrapping Techniques for Polysynthetic Morphological Analysis
William Lane, Steven Bird
Abstract
Polysynthetic languages have exceptionally large and sparse vocabularies, thanks to the number of morpheme slots and combinations in a word. This complexity, together with a general scarcity of written data, poses a challenge to the development of natural language technologies. To address this challenge, we offer linguistically-informed approaches for bootstrapping a neural morphological analyzer, and demonstrate its application to Kunwinjku, a polysynthetic Australian language. We generate data from a finite state transducer to train an encoderdecoder model. We improve the model by "hallucinating" missing linguistic structure into the training data, and by resampling from a Zipf distribution to simulate a more natural distribution of morphemes. The best model accounts for all instances of reduplication in the test set and achieves an accuracy of 94.7% overall, a 10 percentage point improvement over the FST baseline. This process demonstrates the feasibility of bootstrapping a neural morph analyzer from minimal resources.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9b2afbf2-1f67-4cad-a850-a3e1fef9f930Cited by top-tier papers3
- Weakly Supervised Word Segmentation for Computational Language DocumentationShu Okabe, Laurent Besacier, François YvonACL 2022 · 5 citations
- Is linguistically-motivated data augmentation worth it?Ray Groshan, Michael Ginn, Alexis PalmerACL 2025
- Understanding Compositional Data Augmentation in Typologically Diverse Morphological InflectionFarhan Samir, Miikka SilfverbergEMNLP 2023
Related papers
- Local Word Discovery for Interactive TranscriptionWilliam Lane, Steven BirdEMNLP 2021
- Minimal Supervision for Morphological InflectionOmer Goldman, Reut TsarfatyEMNLP 2021
- Interactive Word Completion for Plains CreeWilliam Lane, Atticus Harrigan, Antti ArppeACL 2022 · 3 citations
- Hints on the data for language modeling of synthetic languages with transformersRodolfo Zevallos, Núria BelACL 2023 · 2 citations
- Modeling Morphological Typology for Unsupervised Learning of Language MorphologyHongzhi Xu, Jordan Kodner, Mitchell Marcus, Charles YangACL 2020 · 7 citations
