Making Monolingual Sentence Embeddings Multilingual using Knowledge Distillation
Nils Reimers, Iryna Gurevych
摘要
We present an easy and efficient method to extend existing sentence embedding models to new languages. This allows to create multilingual versions from previously monolingual models. The training is based on the idea that a translated sentence should be mapped to the same location in the vector space as the original sentence. We use the original (monolingual) model to generate sentence embeddings for the source language and then train a new system on translated sentences to mimic the original model. Compared to other methods for training multilingual sentence embeddings, this approach has several advantages: It is easy to extend existing models with relatively few samples to new languages, it is easier to ensure desired properties for the vector space, and the hardware requirements for training are lower. We demonstrate the effectiveness of our approach for 50+ languages from various language families. Code to extend sentence embeddings models to more than 400 languages is publicly available. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper130
- Decomposing NeRF for Editing via Feature Field DistillationSosuke Kobayashi, Eiichi Matsumoto, Vincent SitzmannNeurIPS 2022 · 被引用 479 次
- Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue AbilitiesZhifeng Kong, Arushi Goel, Rohan Badlani, Wei Ping 等ICML 2024 · 被引用 207 次
- A Semantic Invariant Robust Watermark for Large Language ModelsAiwei Liu, Leyi Pan, Xuming Hu, Shiao Meng 等ICLR 2024 · 被引用 108 次
- Knowledge-driven Data Construction for Zero-shot Evaluation in Commonsense Question AnsweringKaixin Ma, Filip Ilievski, Jonathan Francis, Yonatan Bisk 等AAAI 2021 · 被引用 100 次
- Controlled Text Generation as Continuous Optimization with Multiple ConstraintsSachin Kumar, Eric Malmi, Aliaksei Severyn, Yulia TsvetkovNeurIPS 2021 · 被引用 91 次
它引用的顶会 Paper3
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- LAReQA: Language-Agnostic Answer Retrieval from a Multilingual PoolUma Roy, Noah Constant, Rami Al-Rfou, Aditya Barua 等EMNLP 2020 · 被引用 39 次
- Language-agnostic BERT Sentence EmbeddingFangxiaoyu Feng, Yinfei Yang, Daniel Cer, Naveen Arivazhagan 等ACL 2022
相关 Paper
- Cross-lingual Sentence Embedding using Multi-Task LearningKoustava Goswami, Sourav Dutta, Haytham Assem, Theodorus Fransen 等EMNLP 2021 · 被引用 9 次
- Language-agnostic Representation from Multilingual Sentence Encoders for Cross-lingual Similarity EstimationNattapong Tiyajamorn, Tomoyuki Kajiwara, Yuki Arase, Makoto OnizukaEMNLP 2021 · 被引用 16 次
- Beyond Contrastive Learning: A Variational Generative Model for Multilingual RetrievalJohn Wieting, Jonathan H. Clark, William W. Cohen, Graham Neubig 等ACL 2023 · 被引用 3 次
- A Bilingual Generative Transformer for Semantic Sentence EmbeddingJohn Wieting, Graham Neubig, Taylor Berg-KirkpatrickEMNLP 2020 · 被引用 4 次
- Modular Sentence Encoders: Separating Language Specialization from Cross-Lingual AlignmentYongxin Huang, Kexin Wang, Goran Glavas, Iryna GurevychACL 2025
