Towards Higher Pareto Frontier in Multilingual Machine Translation
Yi-Chong Huang, Xiaocheng Feng, Xinwei Geng, Baohang Li, Bing Qin
摘要
Multilingual neural machine translation has witnessed remarkable progress in recent years. However, the long-tailed distribution of multilingual corpora poses a challenge of Pareto optimization, i.e., optimizing for some languages may come at the cost of degrading the performance of others.Existing balancing training strategies are equivalent to a series of Pareto optimal solutions, which trade off on a Pareto frontierIn Pareto optimization, Pareto optimal solutions refer to solutions in which none of the objectives can be improved without sacrificing at least one of the other objectives. The set of all Pareto optimal solutions forms a Pareto frontier..In this work, we propose a new training framework, Pareto Mutual Distillation (Pareto-MD), towards pushing the Pareto frontier outwards rather than making trade-offs.Specifically, Pareto-MD collaboratively trains two Pareto optimal solutions that favor different languages and allows them to learn from the strengths of each other via knowledge distillation.Furthermore, we introduce a novel strategy to enable stronger communication between Pareto optimal solutions and broaden the applicability of our approach. Experimental results on the widely-used WMT and TED datasets show that our method significantly pushes the Pareto frontier and outperforms baselines by up to +2.46 BLEUOur code will be released upon acceptance..
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Exploring Intrinsic Language-specific Subspaces in Fine-tuning Multilingual Neural Machine TranslationZhe Cao, Zhi Qu, Hidetaka Kamigaito, Taro WatanabeEMNLP 2024 · 被引用 1 次
- Language Constrained Multimodal Hyper Adapter For Many-to-Many Multimodal SummarizationNayu Liu, Fanglong Yao, Haoran Luo, Yong Yang 等ACL 2025
它引用的顶会 Paper7
- Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep LearningZeyuan Allen-Zhu, Yuanzhi LiICLR 2023 · 被引用 151 次
- Balancing Training for Multilingual Neural Machine TranslationXinyi Wang, Yulia Tsvetkov, Graham NeubigACL 2020 · 被引用 74 次
- Feature Kernel DistillationBobby He, Mete OzayICLR 2022 · 被引用 18 次
- Distributionally Robust Multilingual Machine TranslationChunting Zhou, Daniel Levy, Xian Li, Marjan Ghazvininejad 等EMNLP 2021 · 被引用 14 次
- Uncertainty-Aware Balancing for Multilingual and Multi-Domain Neural Machine Translation TrainingMinghao Wu, Yitong Li, Meng Zhang, Liangyou Li 等EMNLP 2021 · 被引用 10 次
相关 Paper
- Exploring All-In-One Knowledge Distillation Framework for Neural Machine TranslationZhongjian Miao, Wen Zhang, Jinsong Su, Xiang Li 等EMNLP 2023 · 被引用 5 次
- Unifying the Convergences in Multilingual Neural Machine TranslationYi-Chong Huang, Xiaocheng Feng, Xinwei Geng, Bing QinEMNLP 2022 · 被引用 6 次
- Pretrained Bidirectional Distillation for Machine TranslationYimeng Zhuang, Mei TuACL 2023 · 被引用 3 次
- Knowledge Distillation for Multilingual Unsupervised Neural Machine TranslationHaipeng Sun, Rui Wang, Kehai Chen, Masao Utiyama 等ACL 2020 · 被引用 37 次
- On the Pareto Front of Multilingual Neural Machine TranslationLiang Chen, Shuming Ma, Dongdong Zhang, Furu Wei 等NeurIPS 2023 · 被引用 8 次
