SemFace: Pre-training Encoder and Decoder with a Semantic Interface for Neural Machine Translation
Shuo Ren, Long Zhou, Shujie Liu, Furu Wei, Ming Zhou, Shuai Ma
Abstract
While pre-training techniques are working very well in natural language processing, how to pre-train a decoder and effectively leverage it for neural machine translation (NMT) still remains a tricky issue. The main reason is that the cross-attention module between the encoder and decoder cannot be pre-trained, and the combined encoder-decoder model cannot work well in the fine-tuning stage because the inputs of the decoder cross-attention come from unknown encoder outputs. In this paper, we propose a better pre-training method for NMT by defining a semantic interface (SemFace) between the pre-trained encoder and the pre-trained decoder. Specifically, we propose two types of semantic interfaces, including CL-SemFace which regards cross-lingual embeddings as an interface, and VQ-SemFace which employs vector quantized embeddings to constrain the encoder outputs and decoder inputs in the same language-independent space. We conduct massive experiments on six supervised translation pairs and three unsupervised pairs. Experimental results demonstrate that our proposed Sem-Face can effectively connect the pre-trained encoder and decoder, and achieves significant improvement by 3.7 and 1.5 BLEU points on the two tasks respectively compared with previous pre-training-based NMT models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- SpeechUT: Bridging Speech and Text with Hidden-Unit for Encoder-Decoder Based Speech-Text Pre-trainingZiqiang Zhang, Long Zhou, Junyi Ao, Shujie Liu et al.EMNLP 2022 · 38 citations
- Revisiting Commonsense Reasoning in Machine Translation: Training, Evaluation and ChallengeXuebo Liu, Yutong Wang, Derek F. Wong, Runzhe Zhan et al.ACL 2023 · 3 citations
- SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language ProcessingJunyi Ao, Rui Wang, Long Zhou, Chengyi Wang et al.ACL 2022
Builds on9
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Incorporating BERT into Neural Machine TranslationJinhua Zhu, Yingce Xia, Lijun Wu, Di He et al.ICLR 2020 · 391 citations
- Towards Making the Most of BERT in Neural Machine TranslationJiacheng Yang, Mingxuan Wang, Hao Zhou, Chengqi Zhao et al.AAAI 2020 · 164 citations
Related papers
- Universal Conditional Masked Language Pre-training for Neural Machine TranslationPengfei Li, Liangyou Li, Meng Zhang, Minghao Wu et al.ACL 2022 · 32 citations
- Understanding and Improving Sequence-to-Sequence Pretraining for Neural Machine TranslationWenxuan Wang, Wenxiang Jiao, Yongchang Hao, Xing Wang et al.ACL 2022 · 32 citations
- Deep Fusing Pre-trained Models into Neural Machine TranslationRongxiang Weng, Heng Yu, Weihua Luo, Min ZhangAAAI 2022 · 3 citations
- BERT, mBERT, or BiBERT? A Study on Contextualized Embeddings for Neural Machine TranslationHaoran Xu, Benjamin Van Durme, Kenton W. MurrayEMNLP 2021 · 55 citations
- Jointly Masked Sequence-to-Sequence Model for Non-Autoregressive Neural Machine TranslationJunliang Guo, Linli Xu, Enhong ChenACL 2020 · 56 citations
