Towards Higher Pareto Frontier in Multilingual Machine Translation
Yi-Chong Huang, Xiaocheng Feng, Xinwei Geng, Baohang Li, Bing Qin
Abstract
Multilingual neural machine translation has witnessed remarkable progress in recent years. However, the long-tailed distribution of multilingual corpora poses a challenge of Pareto optimization, i.e., optimizing for some languages may come at the cost of degrading the performance of others.Existing balancing training strategies are equivalent to a series of Pareto optimal solutions, which trade off on a Pareto frontierIn Pareto optimization, Pareto optimal solutions refer to solutions in which none of the objectives can be improved without sacrificing at least one of the other objectives. The set of all Pareto optimal solutions forms a Pareto frontier..In this work, we propose a new training framework, Pareto Mutual Distillation (Pareto-MD), towards pushing the Pareto frontier outwards rather than making trade-offs.Specifically, Pareto-MD collaboratively trains two Pareto optimal solutions that favor different languages and allows them to learn from the strengths of each other via knowledge distillation.Furthermore, we introduce a novel strategy to enable stronger communication between Pareto optimal solutions and broaden the applicability of our approach. Experimental results on the widely-used WMT and TED datasets show that our method significantly pushes the Pareto frontier and outperforms baselines by up to +2.46 BLEUOur code will be released upon acceptance..
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 09691530-e071-426d-a6af-b8514de48a13Cited by top-tier papers2
- Exploring Intrinsic Language-specific Subspaces in Fine-tuning Multilingual Neural Machine TranslationZhe Cao, Zhi Qu, Hidetaka Kamigaito, Taro WatanabeEMNLP 2024 · 1 citation
- Language Constrained Multimodal Hyper Adapter For Many-to-Many Multimodal SummarizationNayu Liu, Fanglong Yao, Haoran Luo, Yong Yang et al.ACL 2025
Builds on7
- Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep LearningZeyuan Allen-Zhu, Yuanzhi LiICLR 2023 · 151 citations
- Balancing Training for Multilingual Neural Machine TranslationXinyi Wang, Yulia Tsvetkov, Graham NeubigACL 2020 · 74 citations
- Feature Kernel DistillationBobby He, Mete OzayICLR 2022 · 18 citations
- Distributionally Robust Multilingual Machine TranslationChunting Zhou, Daniel Levy, Xian Li, Marjan Ghazvininejad et al.EMNLP 2021 · 14 citations
- Uncertainty-Aware Balancing for Multilingual and Multi-Domain Neural Machine Translation TrainingMinghao Wu, Yitong Li, Meng Zhang, Liangyou Li et al.EMNLP 2021 · 10 citations
Related papers
- Exploring All-In-One Knowledge Distillation Framework for Neural Machine TranslationZhongjian Miao, Wen Zhang, Jinsong Su, Xiang Li et al.EMNLP 2023 · 5 citations
- Unifying the Convergences in Multilingual Neural Machine TranslationYi-Chong Huang, Xiaocheng Feng, Xinwei Geng, Bing QinEMNLP 2022 · 6 citations
- Pretrained Bidirectional Distillation for Machine TranslationYimeng Zhuang, Mei TuACL 2023 · 3 citations
- Knowledge Distillation for Multilingual Unsupervised Neural Machine TranslationHaipeng Sun, Rui Wang, Kehai Chen, Masao Utiyama et al.ACL 2020 · 37 citations
- On the Pareto Front of Multilingual Neural Machine TranslationLiang Chen, Shuming Ma, Dongdong Zhang, Furu Wei et al.NeurIPS 2023 · 8 citations
