EMMA-X: An EM-like Multilingual Pre-training Algorithm for Cross-lingual Representation Learning
Ping Guo, Xiangpeng Wei, Yue Hu, Baosong Yang, Dayiheng Liu, Fei Huang, Jun Xie
摘要
Expressing universal semantics common to all languages is helpful in understanding the meanings of complex and culture-specific sentences. The research theme underlying this scenario focuses on learning universal representations across languages with the usage of massive parallel corpora. However, due to the sparsity and scarcity of parallel data, there is still a big challenge in learning authentic "universals" for any two languages. In this paper, we propose EMMA-X: an EM-like Multilingual pre-training Algorithm, to learn (X)Cross-lingual universals with the aid of excessive multilingual non-parallel data. EMMA-X unifies the cross-lingual representation learning task and an extra semantic relation prediction task within an EM framework. Both the extra semantic classifier and the cross-lingual sentence encoder approximate the semantic relation of two sentences, and supervise each other until convergence. To evaluate EMMA-X, we conduct experiments on XRETE, a newly introduced benchmark containing 12 widely studied cross-lingual tasks that fully depend on sentence-level representations. Results reveal that EMMA-X achieves state-of-the-art performance. Further geometric analysis of the built representation space with three requirements demonstrates the superiority of EMMA-X over advanced models 2 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- MuRating: A High Quality Data Selecting Approach to Multilingual Large Language Model PretrainingZhixun Chen, Ping Guo, Wenhan Han, Yifan Zhang 等NeurIPS 2025 · 被引用 4 次
- MTLS: Making Texts into Linguistic SymbolsWenlong Fei, Xiaohua Wang, Min Hu, Qingyu Zhang 等EMNLP 2024 · 被引用 1 次
它引用的顶会 Paper31
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 被引用 2,496 次
- XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual GeneralisationJunjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig 等ICML 2020 · 被引用 1,132 次
- Supervised Contrastive Learning for Pre-trained Language Model Fine-tuningBeliz Gunel, Jingfei Du, Alexis Conneau, Veselin StoyanovICLR 2021 · 被引用 595 次
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
相关 Paper
- ERNIE-M: Enhanced Multilingual Representation by Aligning Cross-lingual Semantics with Monolingual CorporaXuan Ouyang, Shuohuan Wang, Chao Pang, Yu Sun 等EMNLP 2021 · 被引用 68 次
- AlignX: Advancing Multilingual Large Language Models with Multilingual Representation AlignmentMengyu Bu, Shaolei Zhang, Zhongjun He, Hua Wu 等EMNLP 2025
- On Learning Universal Representations Across LanguagesXiangpeng Wei, Rongxiang Weng, Yue Hu, Luxi Xing 等ICLR 2021 · 被引用 93 次
- Pre-training Universal Language RepresentationYian Li, Hai ZhaoACL 2021
- English Contrastive Learning Can Learn Universal Cross-lingual Sentence EmbeddingsYau-Shian Wang, Ashley Wu, Graham NeubigEMNLP 2022 · 被引用 18 次
