CSP: Code-Switching Pre-training for Neural Machine Translation
Zhen Yang, Bojie Hu, Ambyera Han, Shen Huang, Qi Ju
Abstract
This paper proposes a new pre-training method, called Code-Switching Pre-training (CSP for short) for Neural Machine Translation (NMT). Unlike traditional pre-training method which randomly masks some fragments of the input sentence, the proposed CSP randomly replaces some words in the source sentence with their translation words in the target language. Specifically, we firstly perform lexicon induction with unsupervised word embedding mapping between the source and target languages, and then randomly replace some words in the input sentence with their translation words according to the extracted translation lexicons. CSP adopts the encoderdecoder framework: its encoder takes the codemixed sentence as input, and its decoder predicts the replaced fragment of the input sentence. In this way, CSP is able to pre-train the NMT model by explicitly making the most of the cross-lingual alignment information extracted from the source and target monolingual corpus. Additionally, we relieve the pretrainfinetune discrepancy caused by the artificial symbols like [mask]. To verify the effectiveness of the proposed method, we conduct extensive experiments on unsupervised and supervised NMT. Experimental results show that CSP achieves significant improvements over baselines without pre-training or with other pre-training methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 42d3154c-b1b6-4a56-8afa-d4683a73abe0Cited by top-tier papers9
- Universal Conditional Masked Language Pre-training for Neural Machine TranslationPengfei Li, Liangyou Li, Meng Zhang, Minghao Wu et al.ACL 2022 · 32 citations
- ESM All-Atom: Multi-Scale Protein Language Model for Unified Molecular ModelingKangjie Zheng, Siyu Long, Tianyu Lu, Junwei Yang et al.ICML 2024 · 17 citations
- Flow-Adapter Architecture for Unsupervised Machine TranslationYihong Liu, Haris Jabbar, Hinrich SchützeACL 2022 · 9 citations
- Cultural Concept Adaptation on Multimodal ReasoningZhi Li, Yin ZhangEMNLP 2023 · 6 citations
- Cross-Align: Modeling Deep Cross-lingual Interactions for Word AlignmentSiyu Lai, Zhen Yang, Fandong Meng, Yufeng Chen et al.EMNLP 2022 · 6 citations
Builds on2
Related papers
- Alternating Language Modeling for Cross-Lingual Pre-TrainingJian Yang, Shuming Ma, Dongdong Zhang, Shuangzhi Wu et al.AAAI 2020 · 94 citations
- On-the-fly Cross-lingual Masking for Multilingual Pre-trainingXi Ai, Bin FangACL 2023 · 1 citation
- From Machine Translation to Code-Switching: Generating High-Quality Code-Switched TextIshan Tarunesh, Syamantak Kumar, Preethi JyothiACL 2021
- Pre-training Multilingual Neural Machine Translation by Leveraging Alignment InformationZehui Lin, Xiao Pan, Mingxuan Wang, Xipeng Qiu et al.EMNLP 2020 · 82 citations
- Deep Fusing Pre-trained Models into Neural Machine TranslationRongxiang Weng, Heng Yu, Weihua Luo, Min ZhangAAAI 2022 · 3 citations
