Self-supervised and Supervised Joint Training for Resource-rich Machine Translation
Yong Cheng, Wei Wang, Lu Jiang, Wolfgang Macherey
Abstract
Self-supervised pre-training of text representations has been successfully applied to low-resource Neural Machine Translation (NMT). However, it usually fails to achieve notable gains on resource-rich NMT. In this paper, we propose a joint training approach, -XEnDec, to combine self-supervised and supervised learning to optimize NMT models. To exploit complementary self-supervised signals for supervised learning, NMT models are trained on examples that are interbred from monolingual and parallel sentences through a new process called crossover encoder-decoder. Experiments on two resource-rich translation benchmarks, WMT'14 English-German and WMT'14 English-French, demonstrate that our approach achieves substantial improvements over several strong baseline methods and obtains a new state of the art of 46.19 BLEU on English-French when incorporating back translation. Results also show that our approach is capable of improving model robustness to input perturbations such as code-switching noise which frequently appears on social media.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 595d796c-a9ad-4e37-bba8-008e0cab07cdCited by top-tier papers5
- Spatial-Temporal Hypergraph Self-Supervised Learning for Crime PredictionZhonghang Li, Chao Huang, Lianghao Xia, Yong Xu et al.ICDE 2022 · 82 citations
- Learning to Generalize to More: Continuous Semantic Augmentation for Neural Machine TranslationXiangpeng Wei, Heng Yu, Yue Hu, Rongxiang Weng et al.ACL 2022 · 26 citations
- ChA-MAEViT: Unifying Channel-Aware Masked Autoencoders and Multi-Channel Vision Transformers for Improved Cross-Channel LearningChau Pham, Juan C. Caicedo, Bryan A. PlummerNeurIPS 2025 · 11 citations
- GATITOS: Using a New Multilingual Lexicon for Low-resource Machine TranslationAlexander Jones, Isaac Caswell, Orhan Firat, Ishank SaxenaEMNLP 2023 · 3 citations
- Multilingual Mix: Example Interpolation Improves Multilingual Neural Machine TranslationYong Cheng, Ankur Bapna, Orhan Firat, Yuan Cao et al.ACL 2022
Builds on7
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh et al.ICCV 2019 · 5,843 citations
- AugMix: A Simple Data Processing Method to Improve Robustness and UncertaintyDan Hendrycks, Norman Mu, Ekin Dogus Cubuk, Barret Zoph et al.ICLR 2020 · 1,572 citations
- Incorporating BERT into Neural Machine TranslationJinhua Zhu, Yingce Xia, Lijun Wu, Di He et al.ICLR 2020 · 391 citations
- Beyond Synthetic Noise: Deep Learning on Controlled Noisy LabelsLu Jiang, Di Huang, Mason Liu, Weilong YangICML 2020 · 241 citations
- Towards Making the Most of BERT in Neural Machine TranslationJiacheng Yang, Mingxuan Wang, Hao Zhou, Chengqi Zhao et al.AAAI 2020 · 164 citations
Related papers
- Universal Conditional Masked Language Pre-training for Neural Machine TranslationPengfei Li, Liangyou Li, Meng Zhang, Minghao Wu et al.ACL 2022 · 32 citations
- Data Diversification: A Simple Strategy For Neural Machine TranslationXuan-Phi Nguyen, Shafiq R. Joty, Kui Wu, Ai Ti AwNeurIPS 2020 · 75 citations
- CSP: Code-Switching Pre-training for Neural Machine TranslationZhen Yang, Bojie Hu, Ambyera Han, Shen Huang et al.EMNLP 2020 · 62 citations
- Unified Speech-Text Pre-training for Speech Translation and RecognitionYun Tang, Hongyu Gong, Ning Dong, Changhan Wang et al.ACL 2022 · 104 citations
- Understanding and Improving Sequence-to-Sequence Pretraining for Neural Machine TranslationWenxuan Wang, Wenxiang Jiao, Yongchang Hao, Xing Wang et al.ACL 2022 · 32 citations
