Language-agnostic Representation from Multilingual Sentence Encoders for Cross-lingual Similarity Estimation
Nattapong Tiyajamorn, Tomoyuki Kajiwara, Yuki Arase, Makoto Onizuka
Abstract
We propose a method to distil languageagnostic meaning embedding using a multilingual sentence encoder. By removing languagespecific information from the original embedding, we retrieve an embedding that fully represents the meaning of the sentence. The proposed method relies only on parallel corpora without any human annotations. Our meaning embedding allows for efficient cross-lingual sentence similarity estimation using a simple cosine similarity calculation. Experimental results of both the quality estimation of machine translation and cross-lingual semantic textual similarity tasks reveal that our method consistently outperforms the strong baselines using the original multilingual embeddings. The method also consistently improves the performance of any pre-trained multilingual sentence encoder, even in low-resource language pairs, where only tens of thousands of parallel sentence pairs are available. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e4754b68-4c98-4f6f-9203-c8bcd502e082Cited by top-tier papers5
- EMMA-X: An EM-like Multilingual Pre-training Algorithm for Cross-lingual Representation LearningPing Guo, Xiangpeng Wei, Yue Hu, Baosong Yang et al.NeurIPS 2023 · 8 citations
- Language Concept Erasure for Language-invariant Dense RetrievalZhiqi Huang, Puxuan Yu, Shauli Ravfogel, James AllanEMNLP 2024 · 1 citation
- Analysis of Multi-Source Language Training in Cross-Lingual TransferSeong Hoon Lim, Taejun Yun, Jinhyeon Kim, Jihun Choi et al.ACL 2024 · 1 citation
- Shared Path: Unraveling Memorization in Multilingual LLMs through Language SimilaritiesXiaoyu Luo, Yiyi Chen, Johannes Bjerva, Qiongxiu LiEMNLP 2025 · 1 citation
- Modular Sentence Encoders: Separating Language Specialization from Cross-Lingual AlignmentYongxin Huang, Kexin Wang, Goran Glavas, Iryna GurevychACL 2025
Builds on8
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- Cross-Lingual Ability of Multilingual BERT: An Empirical StudyKarthikeyan K, Zihan Wang, Stephen Mayhew, Dan RothICLR 2020 · 378 citations
- Making Monolingual Sentence Embeddings Multilingual using Knowledge DistillationNils Reimers, Iryna GurevychEMNLP 2020 · 54 citations
Related papers
- Discovering Low-rank Subspaces for Language-agnostic Multilingual RepresentationsZhihui Xie, Handong Zhao, Tong Yu, Shuai LiEMNLP 2022 · 3 citations
- Retrofitting Multilingual Sentence Embeddings with Abstract Meaning RepresentationDeng Cai, Xin Li, Jackie Chun-Sing Ho, Lidong Bing et al.EMNLP 2022 · 4 citations
- Cross-lingual Sentence Embedding using Multi-Task LearningKoustava Goswami, Sourav Dutta, Haytham Assem, Theodorus Fransen et al.EMNLP 2021 · 9 citations
- English Contrastive Learning Can Learn Universal Cross-lingual Sentence EmbeddingsYau-Shian Wang, Ashley Wu, Graham NeubigEMNLP 2022 · 18 citations
- A Bilingual Generative Transformer for Semantic Sentence EmbeddingJohn Wieting, Graham Neubig, Taylor Berg-KirkpatrickEMNLP 2020 · 4 citations
