VECO: Variable and Flexible Cross-lingual Pre-training for Language Understanding and Generation
Fuli Luo, Wei Wang, Jiahao Liu, Yijia Liu, Bin Bi, Songfang Huang, Fei Huang, Luo Si
摘要
Existing work in multilingual pretraining has demonstrated the potential of cross-lingual transferability by training a unified Transformer encoder for multiple languages. However, much of this work only relies on the shared vocabulary and bilingual contexts to encourage the correlation across languages, which is loose and implicit for aligning the contextual representations between languages. In this paper, we plug a cross-attention module into the Transformer encoder to explicitly build the interdependence between languages. It can effectively avoid the degeneration of predicting masked words only conditioned on the context in its own language. More importantly, when fine-tuning on downstream tasks, the cross-attention module can be plugged in or out on-demand, thus naturally benefiting a wider range of cross-lingual tasks, from language understanding to generation. As a result, the proposed cross-lingual model delivers new state-of-the-art results on various cross-lingual understanding tasks of the XTREME benchmark, covering text classification, sequence labeling, question answering, and sentence retrieval. For cross-lingual generation tasks, it also outperforms all existing cross-lingual models and state-of-theart Transformer variants on WMT14 Englishto-German and English-to-French translation datasets, with gains of up to 1∼2 BLEU. 1 * Equal contribution. 1 Code and model are available at https://github. com/alibaba/AliceMind/tree/main/VECO b) Translation Language Modeling (TLM) a) Masked Language Modeling (MLM)
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Enhancing Cross-lingual Transfer by Manifold MixupHuiyun Yang, Huadong Chen, Hao Zhou, Lei LiICLR 2022 · 被引用 49 次
- Improving Neural Cross-Lingual Abstractive Summarization via Employing Optimal Transport Distance for Knowledge DistillationThong Thanh Nguyen, Anh Tuan LuuAAAI 2022 · 被引用 46 次
- EMMA-X: An EM-like Multilingual Pre-training Algorithm for Cross-lingual Representation LearningPing Guo, Xiangpeng Wei, Yue Hu, Baosong Yang 等NeurIPS 2023 · 被引用 8 次
- Cross-Align: Modeling Deep Cross-lingual Interactions for Word AlignmentSiyu Lai, Zhen Yang, Fandong Meng, Yufeng Chen 等EMNLP 2022 · 被引用 6 次
- Just Go Parallel: Improving the Multilingual Capabilities of Large Language ModelsMuhammad Reza Qorib, Junyi Li, Hwee Tou NgACL 2025 · 被引用 5 次
它引用的顶会 Paper12
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- Incorporating BERT into Neural Machine TranslationJinhua Zhu, Yingce Xia, Lijun Wu, Di He 等ICLR 2020 · 被引用 391 次
- XGLUE: A New Benchmark Datasetfor Cross-lingual Pre-training, Understanding and GenerationYaobo Liang, Nan Duan, Yeyun Gong, Ning Wu 等EMNLP 2020 · 被引用 232 次
- Cross-Lingual Natural Language Generation via Pre-TrainingZewen Chi, Li Dong, Furu Wei, Wenhui Wang 等AAAI 2020 · 被引用 142 次
相关 Paper
- Rethinking Embedding Coupling in Pre-trained Language ModelsHyung Won Chung, Thibault Févry, Henry Tsai, Melvin Johnson 等ICLR 2021 · 被引用 11 次
- Soft Language Clustering for Multilingual Model Pre-trainingJiali Zeng, Yufan Jiang, Yongjing Yin, Yi Jing 等ACL 2023 · 被引用 1 次
- On Learning Universal Representations Across LanguagesXiangpeng Wei, Rongxiang Weng, Yue Hu, Luxi Xing 等ICLR 2021 · 被引用 93 次
- XLM-K: Improving Cross-Lingual Language Model Pre-training with Multilingual KnowledgeXiaoze Jiang, Yaobo Liang, Weizhu Chen, Nan DuanAAAI 2022 · 被引用 31 次
- Improving Pretrained Cross-Lingual Language Models via Self-Labeled Word AlignmentZewen Chi, Li Dong, Bo Zheng, Shaohan Huang 等ACL 2021
