Unifying the Convergences in Multilingual Neural Machine Translation
Yi-Chong Huang, Xiaocheng Feng, Xinwei Geng, Bing Qin
Abstract
Although all-in-one-model multilingual neural machine translation (multilingual NMT) has achieved remarkable progress, the convergence inconsistency in the joint training is ignored, i.e.,different language pairs reaching convergence in different epochs. This leads to the trained MNMT model over-fitting lowresource language translations while underfitting high-resource ones. In this paper, we propose a novel training strategy named LSSD (Language-Specific Self-Distillation), which can alleviate the convergence inconsistency and help MNMT models achieve the best performance on each language pair simultaneously. Specifically, LSSD picks up language-specific best checkpoints for each language pair to teach the current model on the fly. Furthermore, we systematically explore three sample-level manipulations of knowledge transferring. Experimental results on three datasets show that LSSD obtains consistent improvements towards all language pairs and achieves the state-of-the-art 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 91ac747a-5f7c-4756-92ac-cf5f010ce43bCited by top-tier papers2
- The Zeno's Paradox of 'Low-Resource' LanguagesHellina Hailu Nigatu, Atnafu Lambebo Tonja, Benjamin Rosman, Thamar Solorio et al.EMNLP 2024 · 10 citations
- Towards Higher Pareto Frontier in Multilingual Machine TranslationYi-Chong Huang, Xiaocheng Feng, Xinwei Geng, Baohang Li et al.ACL 2023 · 8 citations
Builds on9
- Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self DistillationLinfeng Zhang, Jiebo Song, Anni Gao, Jingwei Chen et al.ICCV 2019 · 1,069 citations
- Improving Massively Multilingual Neural Machine Translation and Zero-Shot TranslationBiao Zhang, Philip Williams, Ivan Titov, Rico SennrichACL 2020 · 213 citations
- Balancing Training for Multilingual Neural Machine TranslationXinyi Wang, Yulia Tsvetkov, Graham NeubigACL 2020 · 74 citations
- Multi-task Learning for Multilingual Neural Machine TranslationYiren Wang, ChengXiang Zhai, Hany HassanEMNLP 2020 · 58 citations
- Distributionally Robust Multilingual Machine TranslationChunting Zhou, Daniel Levy, Xian Li, Marjan Ghazvininejad et al.EMNLP 2021 · 14 citations
Related papers
- Learning Language Specific Sub-network for Multilingual Machine TranslationZehui Lin, Liwei Wu, Mingxuan Wang, Lei LiACL 2021
- Pretrained Bidirectional Distillation for Machine TranslationYimeng Zhuang, Mei TuACL 2023 · 3 citations
- Knowledge Distillation for Multilingual Unsupervised Neural Machine TranslationHaipeng Sun, Rui Wang, Kehai Chen, Masao Utiyama et al.ACL 2020 · 37 citations
- Enhancing Multilingual Capabilities of Large Language Models through Self-Distillation from Resource-Rich LanguagesYuanchi Zhang, Yile Wang, Zijun Liu, Shuo Wang et al.ACL 2024
- Exploring All-In-One Knowledge Distillation Framework for Neural Machine TranslationZhongjian Miao, Wen Zhang, Jinsong Su, Xiang Li et al.EMNLP 2023 · 5 citations
