Multilingual Mix: Example Interpolation Improves Multilingual Neural Machine Translation
Yong Cheng, Ankur Bapna, Orhan Firat, Yuan Cao, Pidong Wang, Wolfgang Macherey
Abstract
Multilingual neural machine translation models are trained to maximize the likelihood of a mix of examples drawn from multiple language pairs. The dominant inductive bias applied to these models is a shared vocabulary and a shared set of parameters across languages; the inputs and labels corresponding to examples drawn from different language pairs might still reside in distinct subspaces. In this paper, we introduce multilingual crossover encoder-decoder (mXEncDec) to fuse language pairs at an instance level. Our approach interpolates instances from different language pairs into joint 'crossover examples' in order to encourage sharing input and output spaces across languages. To ensure better fusion of examples in multilingual settings, we propose several techniques to improve example interpolation across dissimilar languages under heavy data imbalance. Experiments on a large-scale WMT multilingual dataset demonstrate that our approach significantly improves quality on English-to-Many, Many-to-English and zero-shot translation tasks (from +0.5 BLEU up to +5.5 BLEU points). Results on code-switching sets demonstrate the capability of our approach to improve model generalization to out-of-distribution multilingual examples. We also conduct qualitative and quantitative representation comparisons to analyze the advantages of our approach at the representation level.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c83640a9-b450-4e18-a54f-aea86b9bd117Cited by top-tier papers3
- Knowledge Transfer in Incremental Learning for Multilingual Neural Machine TranslationKaiyu Huang, Peng Li, Jin Ma, Ting Yao et al.ACL 2023 · 17 citations
- Mu2SLAM: Multitask, Multilingual Speech and Language ModelsYong Cheng, Yu Zhang, Melvin Johnson, Wolfgang Macherey et al.ICML 2023 · 10 citations
- Learn and Consolidate: Continual Adaptation for Zero-Shot and Multilingual Neural Machine TranslationKaiyu Huang, Peng Li, Junpeng Liu, Maosong Sun et al.EMNLP 2023 · 4 citations
Builds on7
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh et al.ICCV 2019 · 5,843 citations
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- DivideMix: Learning with Noisy Labels as Semi-supervised LearningJunnan Li, Richard Socher, Steven C. H. HoiICLR 2020 · 1,326 citations
- Improving Massively Multilingual Neural Machine Translation and Zero-Shot TranslationBiao Zhang, Philip Williams, Ivan Titov, Rico SennrichACL 2020 · 213 citations
- Emerging Cross-lingual Structure in Pretrained Language ModelsAlexis Conneau, Shijie Wu, Haoran Li, Luke Zettlemoyer et al.ACL 2020 · 210 citations
Related papers
- Self-supervised and Supervised Joint Training for Resource-rich Machine TranslationYong Cheng, Wei Wang, Lu Jiang, Wolfgang MachereyICML 2021 · 19 citations
- Learning Language Specific Sub-network for Multilingual Machine TranslationZehui Lin, Liwei Wu, Mingxuan Wang, Lei LiACL 2021
- Decoupled Vocabulary Learning Enables Zero-Shot Translation from Unseen LanguagesCarlos Mullov, Ngoc-Quan Pham, Alexander WaibelACL 2024 · 1 citation
- Knowledge Distillation for Multilingual Unsupervised Neural Machine TranslationHaipeng Sun, Rui Wang, Kehai Chen, Masao Utiyama et al.ACL 2020 · 37 citations
- Cross-Lingual Pre-Training Based Transfer for Zero-Shot Neural Machine TranslationBaijun Ji, Zhirui Zhang, Xiangyu Duan, Min Zhang et al.AAAI 2020 · 67 citations
