Adapting High-resource NMT Models to Translate Low-resource Related Languages without Parallel Data
Wei-Jen Ko, Ahmed El-Kishky, Adithya Renduchintala, Vishrav Chaudhary, Naman Goyal, Francisco Guzmán, Pascale Fung, Philipp Koehn, Mona T. Diab
摘要
The scarcity of parallel data is a major obstacle for training high-quality machine translation systems for low-resource languages. Fortunately, some low-resource languages are linguistically related or similar to high-resource languages; these related languages may share many lexical or syntactic structures. In this work, we exploit this linguistic overlap to facilitate translating to and from a lowresource language with only monolingual data, in addition to any parallel data in the related high-resource language. Our method, NMT-Adapt, combines denoising autoencoding, back-translation and adversarial objectives to utilize monolingual data for lowresource adaptation. We experiment on 7 languages from three different language families and show that our technique significantly improves translation into low-resource language compared to other translation baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- CogTaskonomy: Cognitively Inspired Task Taxonomy Is Beneficial to Transfer Learning in NLPYifei Luo, Minghui Xu, Deyi XiongACL 2022 · 被引用 20 次
- Continual Learning of Neural Machine Translation within Low Forgetting Risk RegionsShuhao Gu, Bojie Hu, Yang FengEMNLP 2022 · 被引用 11 次
- Understanding In-Context Machine Translation for Low-Resource Languages: A Case Study on ManchuRenhao Pei, Yihong Liu, Peiqin Lin, François Yvon 等ACL 2025 · 被引用 11 次
- The Zeno's Paradox of 'Low-Resource' LanguagesHellina Hailu Nigatu, Atnafu Lambebo Tonja, Benjamin Rosman, Thamar Solorio 等EMNLP 2024 · 被引用 10 次
- Don't Go Far Off: An Empirical Study on Neural Poetry TranslationTuhin Chakrabarty, Arkadiy Saakyan, Smaranda MuresanEMNLP 2021 · 被引用 8 次
它引用的顶会 Paper4
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- Cross-lingual Retrieval for Iterative Self-Supervised TrainingChau Tran, Yuqing Tang, Xian Li, Jiatao GuNeurIPS 2020 · 被引用 76 次
- Unsupervised Neural Dialect Translation with Commonality and Diversity ModelingYu Wan, Baosong Yang, Derek F. Wong, Lidia S. Chao 等AAAI 2020 · 被引用 21 次
- CCMatrix: Mining Billions of High-Quality Parallel Sentences on the WebHolger Schwenk, Guillaume Wenzek, Sergey Edunov, Edouard Grave 等ACL 2021
相关 Paper
- Language Model Prior for Low-Resource Neural Machine TranslationChristos Baziotis, Barry Haddow, Alexandra BirchEMNLP 2020 · 被引用 11 次
- Multilingual Unsupervised Neural Machine Translation with Denoising AdaptersAhmet Üstün, Alexandre Berard, Laurent Besacier, Matthias GalléEMNLP 2021 · 被引用 2 次
- Small Data, Big Impact: Leveraging Minimal Data for Effective Machine TranslationJean Maillard, Cynthia Gao, Elahe Kalbassi, Kaushik Ram Sadagopan 等ACL 2023 · 被引用 3 次
- Iterative Domain-Repaired Back-TranslationHao-Ran Wei, Zhirui Zhang, Boxing Chen, Weihua LuoEMNLP 2020 · 被引用 14 次
- Neural Machine Translation with Monolingual Translation MemoryDeng Cai, Yan Wang, Huayang Li, Wai Lam 等ACL 2021
