Cross-lingual Sentence Embedding using Multi-Task Learning
Koustava Goswami, Sourav Dutta, Haytham Assem, Theodorus Fransen, John P. McCrae
摘要
Multilingual sentence embeddings capture rich semantic information not only for measuring similarity between texts but also for catering to a broad range of downstream crosslingual NLP tasks. State-of-the-art multilingual sentence embedding models require large parallel corpora to learn efficiently, which confines the scope of these models. In this paper, we propose a novel sentence embedding framework based on an unsupervised loss function for generating effective multilingual sentence embeddings, eliminating the need for parallel corpora. We capture semantic similarity and relatedness between sentences using a multitask loss function for training a dual encoder model mapping different languages onto the same vector space. We demonstrate the efficacy of an unsupervised as well as a weakly supervised variant of our framework on STS, BUCC and Tatoeba benchmark tasks. The proposed unsupervised sentence embedding framework outperforms even supervised stateof-the-art methods for certain under-resourced languages on the Tatoeba dataset and on a monolingual benchmark. Further, we show enhanced zero-shot learning capabilities for more than 30 languages, with the model being trained on only 13 languages. Our model can be extended to a wide range of languages from any language family, as it overcomes the requirement of parallel corpora for training.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Non-Linguistic Supervision for Contrastive Learning of Sentence EmbeddingsYiren Jian, Chongyang Gao, Soroush VosoughiNeurIPS 2022 · 被引用 20 次
- English Contrastive Learning Can Learn Universal Cross-lingual Sentence EmbeddingsYau-Shian Wang, Ashley Wu, Graham NeubigEMNLP 2022 · 被引用 18 次
- EMMA-X: An EM-like Multilingual Pre-training Algorithm for Cross-lingual Representation LearningPing Guo, Xiangpeng Wei, Yue Hu, Baosong Yang 等NeurIPS 2023 · 被引用 8 次
- Beyond Contrastive Learning: A Variational Generative Model for Multilingual RetrievalJohn Wieting, Jonathan H. Clark, William W. Cohen, Graham Neubig 等ACL 2023 · 被引用 3 次
- MTLS: Making Texts into Linguistic SymbolsWenlong Fei, Xiaohua Wang, Min Hu, Qingyu Zhang 等EMNLP 2024 · 被引用 1 次
它引用的顶会 Paper6
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- Multilingual Alignment of Contextual Word RepresentationsSteven Cao, Nikita Kitaev, Dan KleinICLR 2020 · 被引用 211 次
- An Unsupervised Sentence Embedding Method by Mutual Information MaximizationYan Zhang, Ruidan He, Zuozhu Liu, Kwan Hui Lim 等EMNLP 2020 · 被引用 126 次
- Making Monolingual Sentence Embeddings Multilingual using Knowledge DistillationNils Reimers, Iryna GurevychEMNLP 2020 · 被引用 54 次
- Emu: Enhancing Multilingual Sentence Embeddings with Semantic SpecializationWataru Hirota, Yoshihiko Suhara, Behzad Golshan, Wang-Chiew TanAAAI 2020 · 被引用 5 次
相关 Paper
- An Ensemble Distillation Framework for Sentence Embeddings with Multilingual Round-Trip TranslationTianyu Zong, Likun ZhangAAAI 2023 · 被引用 1 次
- Language-agnostic BERT Sentence EmbeddingFangxiaoyu Feng, Yinfei Yang, Daniel Cer, Naveen Arivazhagan 等ACL 2022
- Unsupervised Interlingual Semantic Representations from Sentence Embeddings for Zero-Shot Cross-Lingual TransferChanny Hong, Jaeyeon Lee, Jungkwon LeeAAAI 2020 · 被引用 1 次
- Language-agnostic Representation from Multilingual Sentence Encoders for Cross-lingual Similarity EstimationNattapong Tiyajamorn, Tomoyuki Kajiwara, Yuki Arase, Makoto OnizukaEMNLP 2021 · 被引用 16 次
- Modular Sentence Encoders: Separating Language Specialization from Cross-Lingual AlignmentYongxin Huang, Kexin Wang, Goran Glavas, Iryna GurevychACL 2025
