A Latent Morphology Model for Open-Vocabulary Neural Machine Translation
Duygu Ataman, Wilker Aziz, Alexandra Birch
Abstract
Translation into morphologically-rich languages challenges neural machine translation (NMT) models with extremely sparse vocabularies where atomic treatment of surface forms is unrealistic. This problem is typically addressed by either pre-processing words into subword units or performing translation directly at the level of characters. The former is based on word segmentation algorithms optimized using corpus-level statistics with no regard to the translation task. The latter learns directly from translation data but requires rather deep architectures. In this paper, we propose to translate words by modeling word formation through a hierarchical latent variable model which mimics the process of morphological inflection. Our model generates words one character at a time by composing two latent representations: a continuous one, aimed at capturing the lexical semantics, and a set of (approximately) discrete features, aimed at capturing the morphosyntactic function, which are shared among different surface forms. Our model achieves better accuracy in translation into three morphologically-rich languages than conventional open-vocabulary NMT methods, while also demonstrating a better generalization capacity under low to mid-resource settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 55698d4c-0248-42a8-bc0d-c01eeff6834bCited by top-tier papers2
- Evaluating the Morphosyntactic Well-formedness of Generated TextsAdithya Pratapa, Antonios Anastasopoulos, Shruti Rijhwani, Aditi Chaudhary et al.EMNLP 2021 · 6 citations
- Grounded Compositional Outputs for Adaptive Language ModelingNikolaos Pappas, Phoebe Mulcaire, Noah A. SmithEMNLP 2020 · 1 citation
Related papers
- From Characters to Words: Hierarchical Pre-trained Language Model for Open-vocabulary Language UnderstandingLi Sun, Florian Luisier, Kayhan Batmanghelich, Dinei A. F. Florêncio et al.ACL 2023
- Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language ModelsPit Neitemeier, Björn Deiseroth, Constantin Eichenberg, Lukas BallesICLR 2025
- TAMS: Translation-Assisted Morphological SegmentationEnora Rice, Ali Marashian, Luke Gessler, Alexis Palmer et al.ACL 2024
- MorphTE: Injecting Morphology in Tensorized EmbeddingsGuobing Gan, Peng Zhang, Sunzhu Li, Xiuqing Lu et al.NeurIPS 2022 · 11 citations
- Mirror-Generative Neural Machine TranslationZaixiang Zheng, Hao Zhou, Shujian Huang, Lei Li et al.ICLR 2020 · 37 citations
