Multi-task Learning for Multilingual Neural Machine Translation
Yiren Wang, ChengXiang Zhai, Hany Hassan
摘要
While monolingual data has been shown to be useful in improving bilingual neural machine translation (NMT), effectively and efficiently leveraging monolingual data for Multilingual NMT (MNMT) systems is a less explored area. In this work, we propose a multi-task learning (MTL) framework that jointly trains the model with the translation task on bitext data and two denoising tasks on the monolingual data. We conduct extensive empirical studies on MNMT systems with 10 language pairs from WMT datasets. We show that the proposed approach can effectively improve the translation quality for both high-resource and low-resource languages with large margin, achieving significantly better results than the individual bilingual models. We also demonstrate the efficacy of the proposed approach in the zero-shot setup for language pairs without bitext training data. Furthermore, we show the effectiveness of MTL over pre-training approaches for both NMT and cross-lingual transfer learning NLU tasks; the proposed approach outperforms massive scale models trained on single task.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- On the Representation Collapse of Sparse Mixture of ExpertsZewen Chi, Li Dong, Shaohan Huang, Damai Dai 等NeurIPS 2022 · 被引用 223 次
- Gating Dropout: Communication-efficient Regularization for Sparsely Activated TransformersRui Liu, Young Jin Kim, Alexandre Muzio, Hany HassanICML 2022 · 被引用 31 次
- Z-Code++: A Pre-trained Language Model Optimized for Abstractive SummarizationPengcheng He, Baolin Peng, Song Wang, Yang Liu 等ACL 2023 · 被引用 27 次
- BiTIIMT: A Bilingual Text-infilling Method for Interactive Machine TranslationYanling Xiao, Lemao Liu, Guoping Huang, Qu Cui 等ACL 2022 · 被引用 21 次
- Improving Multilingual Translation by Representation and Gradient RegularizationYilin Yang, Akiko Eriguchi, Alexandre Muzio, Prasad Tadepalli 等EMNLP 2021 · 被引用 16 次
它引用的顶会 Paper3
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- XGLUE: A New Benchmark Datasetfor Cross-lingual Pre-training, Understanding and GenerationYaobo Liang, Nan Duan, Yeyun Gong, Ning Wu 等EMNLP 2020 · 被引用 232 次
- Evaluating the Cross-Lingual Effectiveness of Massively Multilingual Neural Machine TranslationAditya Siddhant, Melvin Johnson, Henry Tsai, Naveen Ari 等AAAI 2020 · 被引用 74 次
相关 Paper
- Cross-Lingual Pre-Training Based Transfer for Zero-Shot Neural Machine TranslationBaijun Ji, Zhirui Zhang, Xiangyu Duan, Min Zhang 等AAAI 2020 · 被引用 67 次
- Zero-Shot Cross-Lingual Transfer of Neural Machine Translation with Multilingual Pretrained EncodersGuanhua Chen, Shuming Ma, Yun Chen, Li Dong 等EMNLP 2021 · 被引用 30 次
- Knowledge Distillation for Multilingual Unsupervised Neural Machine TranslationHaipeng Sun, Rui Wang, Kehai Chen, Masao Utiyama 等ACL 2020 · 被引用 37 次
- Neural Machine Translation with Monolingual Translation MemoryDeng Cai, Yan Wang, Huayang Li, Wai Lam 等ACL 2021
- Improving Low-Resource Languages in Pre-Trained Multilingual Language ModelsViktor Hangya, Hossain Shaikh Saadi, Alexander FraserEMNLP 2022 · 被引用 17 次
