Understanding and Improving Sequence-to-Sequence Pretraining for Neural Machine Translation
Wenxuan Wang, Wenxiang Jiao, Yongchang Hao, Xing Wang, Shuming Shi, Zhaopeng Tu, Michael R. Lyu
Abstract
In this paper, we present a substantial step in better understanding the SOTA sequence-to-sequence (Seq2Seq) pretraining for neural machine translation (NMT). We focus on studying the impact of the jointly pretrained decoder, which is the main difference between Seq2Seq pretraining and previous encoder-based pretraining approaches for NMT. By carefully designing experiments on three language pairs, we find that Seq2Seq pretraining is a double-edged sword: On one hand, it helps NMT models to produce more diverse translations and reduce adequacy-related translation errors. On the other hand, the discrepancies between Seq2Seq pretraining and NMT finetuning limit the translation quality (i.e., domain discrepancy) and induce the over-estimation issue (i.e., objective discrepancy). Based on these observations, we further propose simple and effective strategies, named in-domain pretraining and input adaptation to remedy the domain and objective discrepancies, respectively. Experimental results on several language pairs show that our approach can consistently improve both translation performance and model robustness upon Seq2Seq pretraining.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext da6e3657-8806-48b9-9cac-4c12dfdc2cb2Cited by top-tier papers13
- Domain Knowledge Matters: Improving Prompts with Fix Templates for Repairing Python Type ErrorsYun Peng, Shuzheng Gao, Cuiyun Gao, Yintong Huo et al.ICSE 2024 · 39 citations
- CLIPTrans: Transferring Visual Knowledge with Pre-trained Models for Multimodal Machine TranslationDevaansh Gupta, Siddhant Kharbanda, Jiawei Zhou, Wanhua Li et al.ICCV 2023 · 28 citations
- MTTM: Metamorphic Testing for Textual Content Moderation SoftwareWenxuan Wang, Jen-tse Huang, Weibin Wu, Jianping Zhang et al.ICSE 2023 · 23 citations
- AEON: a method for automatic evaluation of NLP test casesJen-tse Huang, Jianping Zhang, Wenxuan Wang, Pinjia He et al.ISSTA 2022 · 18 citations
- Adaptive Few-shot Prompting for Machine Translation with Pre-trained Language ModelsLei Tang, Jinghui Qin, Wenxuan Ye, Hao Tan et al.AAAI 2025 · 9 citations
Builds on11
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- Incorporating BERT into Neural Machine TranslationJinhua Zhu, Yingce Xia, Lijun Wu, Di He et al.ICLR 2020 · 391 citations
- Towards Making the Most of BERT in Neural Machine TranslationJiacheng Yang, Mingxuan Wang, Hao Zhou, Chengqi Zhao et al.AAAI 2020 · 164 citations
- Order-Agnostic Cross Entropy for Non-Autoregressive Machine TranslationCunxiao Du, Zhaopeng Tu, Jing JiangICML 2021 · 93 citations
Related papers
- Deep Fusing Pre-trained Models into Neural Machine TranslationRongxiang Weng, Heng Yu, Weihua Luo, Min ZhangAAAI 2022 · 3 citations
- Universal Conditional Masked Language Pre-training for Neural Machine TranslationPengfei Li, Liangyou Li, Meng Zhang, Minghao Wu et al.ACL 2022 · 32 citations
- Incorporating BERT into Parallel Sequence Decoding with AdaptersJunliang Guo, Zhirui Zhang, Linli Xu, Hao-Ran Wei et al.NeurIPS 2020 · 72 citations
- Jointly Masked Sequence-to-Sequence Model for Non-Autoregressive Neural Machine TranslationJunliang Guo, Linli Xu, Enhong ChenACL 2020 · 56 citations
- Acquiring Knowledge from Pre-Trained Model to Neural Machine TranslationRongxiang Weng, Heng Yu, Shujian Huang, Shanbo Cheng et al.AAAI 2020 · 71 citations
