On the Pareto Front of Multilingual Neural Machine Translation
Liang Chen, Shuming Ma, Dongdong Zhang, Furu Wei, Baobao Chang
摘要
In this work, we study how the performance of a given direction changes with its sampling ratio in Multilingual Neural Machine Translation (MNMT). By training over 200 multilingual models with various model sizes, data sizes, and language directions, we find it interesting that the performance of certain translation direction does not always improve with the increase of its weight in the multi-task optimization objective. Accordingly, scalarization method leads to a multitask trade-off front that deviates from the traditional Pareto front when there exists data imbalance in the training corpus, which poses a great challenge to improve the overall performance of all directions. Based on our observations, we propose the Double Power Law to predict the unique performance trade-off front in MNMT, which is robust across various languages, data adequacy, and the number of tasks. Finally, we formulate the sample ratio selection problem in MNMT as an optimization problem based on the Double Power Law. In our experiments, it achieves better performance than temperature searching and gradient manipulation methods with only 1/5 to 1/2 of the total training budget. We release the code at https://github.com/pkunlp-icler/ParetoMNMT for reproduction.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- MMICL: Empowering Vision-language Model with Multi-Modal In-Context LearningHaozhe Zhao, Zefan Cai, Shuzheng Si, Xiaojian Ma 等ICLR 2024 · 被引用 206 次
- Soul-Mix: Enhancing Multimodal Machine Translation with Manifold MixupXuxin Cheng, Ziyu Yao, Yifei Xin, Hao An 等ACL 2024 · 被引用 3 次
- Exploring Polyglot Harmony: On Multilingual Data Allocation for Large Language Models PretrainingPing Guo, Yubing Ren, Binbin Liu, Fengze Liu 等NeurIPS 2025 · 被引用 2 次
- Exploring Intrinsic Language-specific Subspaces in Fine-tuning Multilingual Neural Machine TranslationZhe Cao, Zhi Qu, Hidetaka Kamigaito, Taro WatanabeEMNLP 2024 · 被引用 1 次
它引用的顶会 Paper11
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine 等NeurIPS 2020 · 被引用 2,261 次
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 被引用 1,861 次
- Gradient Vaccine: Investigating and Improving Multi-task Optimization in Massively Multilingual ModelsZirui Wang, Yulia Tsvetkov, Orhan Firat, Yuan CaoICLR 2021 · 被引用 241 次
- Do Current Multi-Task Optimization Methods in Deep Learning Even Help?Derrick Xin, Behrooz Ghorbani, Justin Gilmer, Ankush Garg 等NeurIPS 2022 · 被引用 91 次
相关 Paper
- Scaling Laws for Multilingual Neural Machine TranslationPatrick Fernandes, Behrooz Ghorbani, Xavier Garcia, Markus Freitag 等ICML 2023 · 被引用 37 次
- Balancing Training for Multilingual Neural Machine TranslationXinyi Wang, Yulia Tsvetkov, Graham NeubigACL 2020 · 被引用 74 次
- Distributionally Robust Multilingual Machine TranslationChunting Zhou, Daniel Levy, Xian Li, Marjan Ghazvininejad 等EMNLP 2021 · 被引用 14 次
- Towards Higher Pareto Frontier in Multilingual Machine TranslationYi-Chong Huang, Xiaocheng Feng, Xinwei Geng, Baohang Li 等ACL 2023 · 被引用 8 次
- Causes and Cures for Interference in Multilingual TranslationUri Shaham, Maha Elbayad, Vedanuj Goswami, Omer Levy 等ACL 2023 · 被引用 6 次
