XLM-E: Cross-lingual Language Model Pre-training via ELECTRA
Zewen Chi, Shaohan Huang, Li Dong, Shuming Ma, Bo Zheng, Saksham Singhal, Payal Bajaj, Xia Song, Xian-Ling Mao, Heyan Huang, Furu Wei
Abstract
In this paper, we introduce ELECTRA-style tasks (Clark et al., 2020b) to cross-lingual language model pre-training. Specifically, we present two pre-training tasks, namely multilingual replaced token detection, and translation replaced token detection. Besides, we pretrain the model, named as XLM-E, on both multilingual and parallel corpora. Our model outperforms the baseline models on various cross-lingual understanding tasks with much less computation cost. Moreover, analysis shows that XLM-E tends to obtain better cross-lingual transferability. * Equal contribution. Zewen Chi contributes during internship at Microsoft Research.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2f5c8f77-824b-45fd-a1ba-45d64ed093d5Cited by top-tier papers29
- VLMo: Unified Vision-Language Pre-Training with Mixture-of-Modality-ExpertsHangbo Bao, Wenhui Wang, Li Dong, Qiang Liu et al.NeurIPS 2022 · 790 citations
- BEATs: Audio Pre-Training with Acoustic TokenizersSanyuan Chen, Yu Wu, Chengyi Wang, Shujie Liu et al.ICML 2023 · 568 citations
- FLAVA: A Foundational Language And Vision Alignment ModelAmanpreet Singh, Ronghang Hu, Vedanuj Goswami, Guillaume Couairon et al.CVPR 2022 · 483 citations
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding SharingPengcheng He, Jianfeng Gao, Weizhu ChenICLR 2023 · 394 citations
- Corrupted Image Modeling for Self-Supervised Visual Pre-TrainingYuxin Fang, Li Dong, Hangbo Bao, Xinggang Wang et al.ICLR 2023 · 23 citations
Builds on12
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 541 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- UniLMv2: Pseudo-Masked Language Models for Unified Language Model Pre-TrainingHangbo Bao, Li Dong, Furu Wei, Wenhui Wang et al.ICML 2020 · 423 citations
- COCO-LM: Correcting and Contrasting Text Sequences for Language Model PretrainingYu Meng, Chenyan Xiong, Payal Bajaj, Saurabh Tiwary et al.NeurIPS 2021 · 231 citations
- Multilingual Alignment of Contextual Word RepresentationsSteven Cao, Nikita Kitaev, Dan KleinICLR 2020 · 211 citations
Related papers
- Improving Pretrained Cross-Lingual Language Models via Self-Labeled Word AlignmentZewen Chi, Li Dong, Bo Zheng, Shaohan Huang et al.ACL 2021
- XLM-K: Improving Cross-Lingual Language Model Pre-training with Multilingual KnowledgeXiaoze Jiang, Yaobo Liang, Weizhu Chen, Nan DuanAAAI 2022 · 31 citations
- VECO: Variable and Flexible Cross-lingual Pre-training for Language Understanding and GenerationFuli Luo, Wei Wang, Jiahao Liu, Yijia Liu et al.ACL 2021
- Cross-lingual Language Model Pretraining for RetrievalPuxuan Yu, Hongliang Fei, Ping LiWWW 2021 · 42 citations
- Soft Language Clustering for Multilingual Model Pre-trainingJiali Zeng, Yufan Jiang, Yongjing Yin, Yi Jing et al.ACL 2023 · 1 citation
