Cross-lingual Sentence Embedding using Multi-Task Learning
Koustava Goswami, Sourav Dutta, Haytham Assem, Theodorus Fransen, John P. McCrae
Abstract
Multilingual sentence embeddings capture rich semantic information not only for measuring similarity between texts but also for catering to a broad range of downstream crosslingual NLP tasks. State-of-the-art multilingual sentence embedding models require large parallel corpora to learn efficiently, which confines the scope of these models. In this paper, we propose a novel sentence embedding framework based on an unsupervised loss function for generating effective multilingual sentence embeddings, eliminating the need for parallel corpora. We capture semantic similarity and relatedness between sentences using a multitask loss function for training a dual encoder model mapping different languages onto the same vector space. We demonstrate the efficacy of an unsupervised as well as a weakly supervised variant of our framework on STS, BUCC and Tatoeba benchmark tasks. The proposed unsupervised sentence embedding framework outperforms even supervised stateof-the-art methods for certain under-resourced languages on the Tatoeba dataset and on a monolingual benchmark. Further, we show enhanced zero-shot learning capabilities for more than 30 languages, with the model being trained on only 13 languages. Our model can be extended to a wide range of languages from any language family, as it overcomes the requirement of parallel corpora for training.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 540ca2fd-a49f-4b82-bdad-7a84655958acCited by top-tier papers6
- Non-Linguistic Supervision for Contrastive Learning of Sentence EmbeddingsYiren Jian, Chongyang Gao, Soroush VosoughiNeurIPS 2022 · 20 citations
- English Contrastive Learning Can Learn Universal Cross-lingual Sentence EmbeddingsYau-Shian Wang, Ashley Wu, Graham NeubigEMNLP 2022 · 18 citations
- EMMA-X: An EM-like Multilingual Pre-training Algorithm for Cross-lingual Representation LearningPing Guo, Xiangpeng Wei, Yue Hu, Baosong Yang et al.NeurIPS 2023 · 8 citations
- Beyond Contrastive Learning: A Variational Generative Model for Multilingual RetrievalJohn Wieting, Jonathan H. Clark, William W. Cohen, Graham Neubig et al.ACL 2023 · 3 citations
- MTLS: Making Texts into Linguistic SymbolsWenlong Fei, Xiaohua Wang, Min Hu, Qingyu Zhang et al.EMNLP 2024 · 1 citation
Builds on6
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- Multilingual Alignment of Contextual Word RepresentationsSteven Cao, Nikita Kitaev, Dan KleinICLR 2020 · 211 citations
- An Unsupervised Sentence Embedding Method by Mutual Information MaximizationYan Zhang, Ruidan He, Zuozhu Liu, Kwan Hui Lim et al.EMNLP 2020 · 126 citations
- Making Monolingual Sentence Embeddings Multilingual using Knowledge DistillationNils Reimers, Iryna GurevychEMNLP 2020 · 54 citations
- Emu: Enhancing Multilingual Sentence Embeddings with Semantic SpecializationWataru Hirota, Yoshihiko Suhara, Behzad Golshan, Wang-Chiew TanAAAI 2020 · 5 citations
Related papers
- An Ensemble Distillation Framework for Sentence Embeddings with Multilingual Round-Trip TranslationTianyu Zong, Likun ZhangAAAI 2023 · 1 citation
- Language-agnostic BERT Sentence EmbeddingFangxiaoyu Feng, Yinfei Yang, Daniel Cer, Naveen Arivazhagan et al.ACL 2022
- Unsupervised Interlingual Semantic Representations from Sentence Embeddings for Zero-Shot Cross-Lingual TransferChanny Hong, Jaeyeon Lee, Jungkwon LeeAAAI 2020 · 1 citation
- Language-agnostic Representation from Multilingual Sentence Encoders for Cross-lingual Similarity EstimationNattapong Tiyajamorn, Tomoyuki Kajiwara, Yuki Arase, Makoto OnizukaEMNLP 2021 · 16 citations
- Modular Sentence Encoders: Separating Language Specialization from Cross-Lingual AlignmentYongxin Huang, Kexin Wang, Goran Glavas, Iryna GurevychACL 2025
