Tackling the Low-resource Challenge for Canonical Segmentation
Manuel Mager, Özlem Çetinoglu, Katharina Kann
Abstract
Canonical morphological segmentation consists of dividing words into their standardized morphemes. Here, we are interested in approaches for the task when training data is limited. We compare model performance in a simulated low-resource setting for the highresource languages German, English, and Indonesian to experiments on new datasets for the truly low-resource languages Popoluca and Tepehua. We explore two new models for the task, borrowing from the closely related area of morphological generation: an LSTM pointer-generator and a sequence-to-sequence model with hard monotonic attention trained with imitation learning. We find that, in the low-resource setting, the novel approaches outperform existing ones on all languages by up to 11.4% accuracy. However, while accuracy in emulated low-resource scenarios is over 50% for all languages, for the truly lowresource languages Popoluca and Tepehua, our best model only obtains 37.4% and 28.4% accuracy, respectively. Thus, we conclude that canonical segmentation is still a challenging task for low-resource languages.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- The Zeno's Paradox of 'Low-Resource' LanguagesHellina Hailu Nigatu, Atnafu Lambebo Tonja, Benjamin Rosman, Thamar Solorio et al.EMNLP 2024 · 10 citations
- Massively Multilingual Joint Segmentation and GlossingMichael Ginn, Lindia Tjuatja, Enora Rice, Ali Marashian et al.ACL 2026 · 2 citations
- Hints on the data for language modeling of synthetic languages with transformersRodolfo Zevallos, Núria BelACL 2023 · 2 citations
- TAMS: Translation-Assisted Morphological SegmentationEnora Rice, Ali Marashian, Luke Gessler, Alexis Palmer et al.ACL 2024
- To POS Tag or Not to POS Tag: The Impact of POS Tags on Morphological Learning in Low-Resource SettingsSarah R. Moeller, Ling Liu, Mans HuldenACL 2021
Related papers
- Multilingual unsupervised sequence segmentation transfers to extremely low-resource languagesC. M. Downey, Shannon Drizin, Levon Haroutunian, Shivin ThukralACL 2022
- Weakly Supervised Word Segmentation for Computational Language DocumentationShu Okabe, Laurent Besacier, François YvonACL 2022 · 5 citations
- BERT-like Models for Slavic Morpheme SegmentationDmitry Morozov, Lizaveta Astapenka, Anna V. Glazkova, Timur Garipov et al.ACL 2025 · 1 citation
- Is linguistically-motivated data augmentation worth it?Ray Groshan, Michael Ginn, Alexis PalmerACL 2025
- Exploring morphology-aware tokenization: A case study on Spanish language modelingAlba Táboas García, Piotr Przybyla, Leo WannerEMNLP 2025 · 1 citation
