TAMS: Translation-Assisted Morphological Segmentation
Enora Rice, Ali Marashian, Luke Gessler, Alexis Palmer, Katharina von der Wense
摘要
Canonical morphological segmentation is the process of analyzing words into the standard (aka underlying) forms of their constituent morphemes. This is a core task in endangered language documentation, and NLP systems have the potential to dramatically speed up this process. In typical language documentation settings, training data for canonical morpheme segmentation is scarce, making it difficult to train high quality models. However, translation data is often much more abundant, and, in this work, we present a method that attempts to leverage translation data in the canonical segmentation task. We propose a character-level sequence-to-sequence model that incorporates representations of translations obtained from pretrained high-resource monolingual language models as an additional signal. Our model outperforms the baseline in a super-low resource setting but yields mixed results on training splits with more data. Additionally, we find that we can achieve strong performance even without needing difficult-to-obtain word level alignments. While further work is needed to make translations useful in higher-resource settings, our model shows promise in severely resource-constrained settings.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper2
相关 Paper
- Getting The Most Out of Your Training Data: Exploring Unsupervised Tasks for Morphological InflectionAbhishek Purushothama, Adam Wiemerslage, Katharina von der WenseEMNLP 2024
- Massively Multilingual Joint Segmentation and GlossingMichael Ginn, Lindia Tjuatja, Enora Rice, Ali Marashian 等ACL 2026 · 被引用 2 次
- Improving Low-Resource Morphological Inflection via Self-Supervised ObjectivesAdam Wiemerslage, Katharina von der WenseACL 2025 · 被引用 2 次
- A Latent Morphology Model for Open-Vocabulary Neural Machine TranslationDuygu Ataman, Wilker Aziz, Alexandra BirchICLR 2020 · 被引用 18 次
- Exploring morphology-aware tokenization: A case study on Spanish language modelingAlba Táboas García, Piotr Przybyla, Leo WannerEMNLP 2025 · 被引用 1 次
