Improving Multilingual Translation by Representation and Gradient Regularization
Yilin Yang, Akiko Eriguchi, Alexandre Muzio, Prasad Tadepalli, Stefan Lee, Hany Hassan
Abstract
Multilingual Neural Machine Translation (NMT) enables one model to serve all translation directions, including ones that are unseen during training, i.e. zero-shot translation. Despite being theoretically attractive, current models often produce low quality translations -commonly failing to even produce outputs in the right target language. In this work, we observe that off-target translation is dominant even in strong multilingual systems, trained on massive multilingual corpora. To address this issue, we propose a joint approach to regularize NMT models at both representation-level and gradient-level. At the representation level, we leverage an auxiliary target language prediction task to regularize decoder outputs to retain information about the target language. At the gradient level, we leverage a small amount of direct data (in thousands of sentence pairs) to regularize model gradients. Our results demonstrate that our approach is highly effective in both reducing off-target translation occurrences and improving zero-shot translation performance by +5.59 and +10.38 BLEU on WMT and OPUS datasets respectively. Moreover, experiments show that our method also works well when the small amount of direct data is not available. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fb70bcf8-4722-4cc2-b59f-d988d02392deCited by top-tier papers7
- Sparse MoE with Language Guided Routing for Multilingual Machine TranslationXinyu Zhao, Xuxi Chen, Yu Cheng, Tianlong ChenICLR 2024 · 19 citations
- Error Norm Truncation: Robust Training in the Presence of Data Noise for Text Generation ModelsTianjian Li, Haoran Xu, Philipp Koehn, Daniel Khashabi et al.ICLR 2024 · 6 citations
- Learn and Consolidate: Continual Adaptation for Zero-Shot and Multilingual Neural Machine TranslationKaiyu Huang, Peng Li, Junpeng Liu, Maosong Sun et al.EMNLP 2023 · 4 citations
- Towards a Better Understanding of Variations in Zero-Shot Neural Machine Translation PerformanceShaomu Tan, Christof MonzEMNLP 2023 · 2 citations
- THOR-MoE: Hierarchical Task-Guided and Context-Responsive Routing for Neural Machine TranslationYunlong Liang, Fandong Meng, Jie ZhouACL 2025 · 1 citation
Builds on6
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine et al.NeurIPS 2020 · 2,261 citations
- Gradient Vaccine: Investigating and Improving Multi-task Optimization in Massively Multilingual ModelsZirui Wang, Yulia Tsvetkov, Orhan Firat, Yuan CaoICLR 2021 · 241 citations
- Improving Massively Multilingual Neural Machine Translation and Zero-Shot TranslationBiao Zhang, Philip Williams, Ivan Titov, Rico SennrichACL 2020 · 213 citations
- Balancing Training for Multilingual Neural Machine TranslationXinyi Wang, Yulia Tsvetkov, Graham NeubigACL 2020 · 74 citations
- Multi-task Learning for Multilingual Neural Machine TranslationYiren Wang, ChengXiang Zhai, Hany HassanEMNLP 2020 · 58 citations
Related papers
- Improving Zero-Shot Translation by Disentangling Positional InformationDanni Liu, Jan Niehues, James Cross, Francisco Guzmán et al.ACL 2021
- Decoupled Vocabulary Learning Enables Zero-Shot Translation from Unseen LanguagesCarlos Mullov, Ngoc-Quan Pham, Alexander WaibelACL 2024 · 1 citation
- Language Model Prior for Low-Resource Neural Machine TranslationChristos Baziotis, Barry Haddow, Alexandra BirchEMNLP 2020 · 11 citations
- Multi Task Learning For Zero Shot Performance Prediction of Multilingual ModelsKabir Ahuja, Shanu Kumar, Sandipan Dandapat, Monojit ChoudhuryACL 2022
- Multilingual Mix: Example Interpolation Improves Multilingual Neural Machine TranslationYong Cheng, Ankur Bapna, Orhan Firat, Yuan Cao et al.ACL 2022
