Adapting High-resource NMT Models to Translate Low-resource Related Languages without Parallel Data
Wei-Jen Ko, Ahmed El-Kishky, Adithya Renduchintala, Vishrav Chaudhary, Naman Goyal, Francisco Guzmán, Pascale Fung, Philipp Koehn, Mona T. Diab
Abstract
The scarcity of parallel data is a major obstacle for training high-quality machine translation systems for low-resource languages. Fortunately, some low-resource languages are linguistically related or similar to high-resource languages; these related languages may share many lexical or syntactic structures. In this work, we exploit this linguistic overlap to facilitate translating to and from a lowresource language with only monolingual data, in addition to any parallel data in the related high-resource language. Our method, NMT-Adapt, combines denoising autoencoding, back-translation and adversarial objectives to utilize monolingual data for lowresource adaptation. We experiment on 7 languages from three different language families and show that our technique significantly improves translation into low-resource language compared to other translation baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d4ba4342-fe0b-439d-977e-cabbf696465aCited by top-tier papers12
- CogTaskonomy: Cognitively Inspired Task Taxonomy Is Beneficial to Transfer Learning in NLPYifei Luo, Minghui Xu, Deyi XiongACL 2022 · 20 citations
- Continual Learning of Neural Machine Translation within Low Forgetting Risk RegionsShuhao Gu, Bojie Hu, Yang FengEMNLP 2022 · 11 citations
- Understanding In-Context Machine Translation for Low-Resource Languages: A Case Study on ManchuRenhao Pei, Yihong Liu, Peiqin Lin, François Yvon et al.ACL 2025 · 11 citations
- The Zeno's Paradox of 'Low-Resource' LanguagesHellina Hailu Nigatu, Atnafu Lambebo Tonja, Benjamin Rosman, Thamar Solorio et al.EMNLP 2024 · 10 citations
- Don't Go Far Off: An Empirical Study on Neural Poetry TranslationTuhin Chakrabarty, Arkadiy Saakyan, Smaranda MuresanEMNLP 2021 · 8 citations
Builds on4
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- Cross-lingual Retrieval for Iterative Self-Supervised TrainingChau Tran, Yuqing Tang, Xian Li, Jiatao GuNeurIPS 2020 · 76 citations
- Unsupervised Neural Dialect Translation with Commonality and Diversity ModelingYu Wan, Baosong Yang, Derek F. Wong, Lidia S. Chao et al.AAAI 2020 · 21 citations
- CCMatrix: Mining Billions of High-Quality Parallel Sentences on the WebHolger Schwenk, Guillaume Wenzek, Sergey Edunov, Edouard Grave et al.ACL 2021
Related papers
- Language Model Prior for Low-Resource Neural Machine TranslationChristos Baziotis, Barry Haddow, Alexandra BirchEMNLP 2020 · 11 citations
- Multilingual Unsupervised Neural Machine Translation with Denoising AdaptersAhmet Üstün, Alexandre Berard, Laurent Besacier, Matthias GalléEMNLP 2021 · 2 citations
- Small Data, Big Impact: Leveraging Minimal Data for Effective Machine TranslationJean Maillard, Cynthia Gao, Elahe Kalbassi, Kaushik Ram Sadagopan et al.ACL 2023 · 3 citations
- Iterative Domain-Repaired Back-TranslationHao-Ran Wei, Zhirui Zhang, Boxing Chen, Weihua LuoEMNLP 2020 · 14 citations
- Neural Machine Translation with Monolingual Translation MemoryDeng Cai, Yan Wang, Huayang Li, Wai Lam et al.ACL 2021
