Cross-model Back-translated Distillation for Unsupervised Machine Translation
Xuan-Phi Nguyen, Shafiq R. Joty, Thanh-Tung Nguyen, Kui Wu, Ai Ti Aw
摘要
Recent unsupervised machine translation (UMT) systems usually employ three main principles: initialization, language modeling and iterative back-translation, though they may apply them differently. Crucially, iterative back-translation and denoising auto-encoding for language modeling provide data diversity to train the UMT systems. However, the gains from these diversification processes has seemed to plateau. We introduce a novel component to the standard UMT framework called Cross-model Back-translated Distillation (CBD), that is aimed to induce another level of data diversification that existing principles lack. CBD is applicable to all previous UMT approaches. In our experiments, it boosts the performance of the standard UMT methods by 1.5-2.0 BLEU. In particular, in WMT'14 English-French, WMT'16 German-English and English-Romanian, CBD outperforms cross-lingual masked language model (XLM) by 2.3, 2.2 and 1.6 BLEU, respectively. It also yields 1.5--3.3 BLEU improvements in IWSLT English-French and English-German tasks. Through extensive experimental analyses, we show that CBD is effective because it embraces data diversity while other similar variants do not.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Refining Low-Resource Unsupervised Translation by Language Disentanglement of Multilingual Translation ModelXuan-Phi Nguyen, Shafiq R. Joty, Kui Wu, Ai Ti AwNeurIPS 2022 · 被引用 6 次
- Exploring All-In-One Knowledge Distillation Framework for Neural Machine TranslationZhongjian Miao, Wen Zhang, Jinsong Su, Xiang Li 等EMNLP 2023 · 被引用 5 次
- Bridging the Data Gap between Training and Inference for Unsupervised Neural Machine TranslationZhiwei He, Xing Wang, Rui Wang, Shuming Shi 等ACL 2022
- Latent Constraints on Unsupervised Text-Graph Alignment with Information AsymmetryJidong Tian, Wenqing Chen, Yitian Li, Caoyun Fan 等AAAI 2023
它引用的顶会 Paper5
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- Data Diversification: A Simple Strategy For Neural Machine TranslationXuan-Phi Nguyen, Shafiq R. Joty, Kui Wu, Ai Ti AwNeurIPS 2020 · 被引用 75 次
- Mirror-Generative Neural Machine TranslationZaixiang Zheng, Hao Zhou, Shujian Huang, Lei Li 等ICLR 2020 · 被引用 37 次
- On The Evaluation of Machine Translation SystemsTrained With Back-TranslationSergey Edunov, Myle Ott, Marc'Aurelio Ranzato, Michael AuliACL 2020 · 被引用 15 次
相关 Paper
- Knowledge Distillation for Multilingual Unsupervised Neural Machine TranslationHaipeng Sun, Rui Wang, Kehai Chen, Masao Utiyama 等ACL 2020 · 被引用 37 次
- Pretrained Bidirectional Distillation for Machine TranslationYimeng Zhuang, Mei TuACL 2023 · 被引用 3 次
- Understanding and Improving Lexical Choice in Non-Autoregressive TranslationLiang Ding, Longyue Wang, Xuebo Liu, Derek F. Wong 等ICLR 2021 · 被引用 44 次
- Towards Making the Most of Cross-Lingual Transfer for Zero-Shot Neural Machine TranslationGuanhua Chen, Shuming Ma, Yun Chen, Dongdong Zhang 等ACL 2022
- Empirical Regularization for Synthetic Sentence Pairs in Unsupervised Neural Machine TranslationXi Ai, Bin FangAAAI 2021 · 被引用 4 次
