Order Matters in the Presence of Dataset Imbalance for Multilingual Learning
Dami Choi, Derrick Xin, Hamid Dadkhahi, Justin Gilmer, Ankush Garg, Orhan Firat, Chih-Kuan Yeh, Andrew M. Dai, Behrooz Ghorbani
摘要
In this paper, we empirically study the optimization dynamics of multi-task learning, particularly focusing on those that govern a collection of tasks with significant data imbalance. We present a simple yet effective method of pre-training on highresource tasks, followed by fine-tuning on a mixture of high/low-resource tasks. We provide a thorough empirical study and analysis of this method's benefits showing that it achieves consistent improvements relative to the performance tradeoff profile of standard static weighting. We analyze under what data regimes this method is applicable and show its improvements empirically in neural machine translation (NMT) and multi-lingual language modeling. * Equal contribution † Work done as a student researcher at Google. 37th Conference on Neural Information Processing Systems (NeurIPS 2023).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- ContraDiff: Planning Towards High Return States via Contrastive LearningYixiang Shan, Zhengbang Zhu, Ting Long, Qifan Liang 等ICLR 2025
- The Geometry of Sequential Learning: Lie-Bracket Prediction of Transfer OrderJohn SweeneyICML 2026
它引用的顶会 Paper7
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- Do Current Multi-Task Optimization Methods in Deep Learning Even Help?Derrick Xin, Behrooz Ghorbani, Justin Gilmer, Ankush Garg 等NeurIPS 2022 · 被引用 91 次
- Balancing Training for Multilingual Neural Machine TranslationXinyi Wang, Yulia Tsvetkov, Graham NeubigACL 2020 · 被引用 74 次
- Multi-task Learning for Multilingual Neural Machine TranslationYiren Wang, ChengXiang Zhai, Hany HassanEMNLP 2020 · 被引用 58 次
- In Neural Machine Translation, What Does Transfer Learning Transfer?Alham Fikri Aji, Nikolay Bogoychev, Kenneth Heafield, Rico SennrichACL 2020 · 被引用 56 次
相关 Paper
- On the Pareto Front of Multilingual Neural Machine TranslationLiang Chen, Shuming Ma, Dongdong Zhang, Furu Wei 等NeurIPS 2023 · 被引用 8 次
- PiKE: Adaptive Data Mixing for Large-Scale Multi-Task Learning Under Low Gradient ConflictsZeman Li, Yuan Deng, Peilin Zhong, Meisam Razaviyayn 等NeurIPS 2025 · 被引用 8 次
- Distributionally Robust Multilingual Machine TranslationChunting Zhou, Daniel Levy, Xian Li, Marjan Ghazvininejad 等EMNLP 2021 · 被引用 14 次
- Deep Fusing Pre-trained Models into Neural Machine TranslationRongxiang Weng, Heng Yu, Weihua Luo, Min ZhangAAAI 2022 · 被引用 3 次
- Dynamic Data Selection and Weighting for Iterative Back-TranslationZi-Yi Dou, Antonios Anastasopoulos, Graham NeubigEMNLP 2020 · 被引用 46 次
