Balancing Training for Multilingual Neural Machine Translation
Xinyi Wang, Yulia Tsvetkov, Graham Neubig
Abstract
When training multilingual machine translation (MT) models that can translate to/from multiple languages, we are faced with imbalanced training sets: some languages have much more training data than others. Standard practice is to up-sample less resourced languages to increase representation, and the degree of up-sampling has a large effect on the overall performance. In this paper, we propose a method that instead automatically learns how to weight training data through a data scorer that is optimized to maximize performance on all test languages. Experiments on two sets of languages under both one-to-many and manyto-one MT settings show our method not only consistently outperforms heuristic baselines in terms of average performance, but also offers flexible control over the performance of which languages are optimized. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 21f40bf3-c338-4caa-928d-057f51b2143eCited by top-tier papers37
- Gradient Vaccine: Investigating and Improving Multi-task Optimization in Massively Multilingual ModelsZirui Wang, Yulia Tsvetkov, Orhan Firat, Yuan CaoICLR 2021 · 241 citations
- Uni-Perceiver-MoE: Learning Sparse Generalist Models with Conditional MoEsJinguo Zhu, Xizhou Zhu, Wenhai Wang, Xiaohua Wang et al.NeurIPS 2022 · 95 citations
- On Negative Interference in Multilingual Models: Findings and A Meta-Learning TreatmentZirui Wang, Zachary C. Lipton, Yulia TsvetkovEMNLP 2020 · 72 citations
- Turning Tables: Generating Examples from Semi-structured Tables for Endowing Language Models with Reasoning SkillsOri Yoran, Alon Talmor, Jonathan BerantACL 2022 · 57 citations
- Class-Weighted Classification: Trade-offs and Robust ApproachesZiyu Xu, Chen Dan, Justin Khim, Pradeep RavikumarICML 2020 · 53 citations
Builds on1
Related papers
- Distributionally Robust Multilingual Machine TranslationChunting Zhou, Daniel Levy, Xian Li, Marjan Ghazvininejad et al.EMNLP 2021 · 14 citations
- Uncertainty-Aware Balancing for Multilingual and Multi-Domain Neural Machine Translation TrainingMinghao Wu, Yitong Li, Meng Zhang, Liangyou Li et al.EMNLP 2021 · 10 citations
- On the Pareto Front of Multilingual Neural Machine TranslationLiang Chen, Shuming Ma, Dongdong Zhang, Furu Wei et al.NeurIPS 2023 · 8 citations
- Order Matters in the Presence of Dataset Imbalance for Multilingual LearningDami Choi, Derrick Xin, Hamid Dadkhahi, Justin Gilmer et al.NeurIPS 2023 · 9 citations
- Optimizing Data Usage via Differentiable RewardsXinyi Wang, Hieu Pham, Paul Michel, Antonios Anastasopoulos et al.ICML 2020 · 73 citations
