Modular Sentence Encoders: Separating Language Specialization from Cross-Lingual Alignment
Yongxin Huang, Kexin Wang, Goran Glavas, Iryna Gurevych
Abstract
Multilingual sentence encoders (MSEs) are commonly obtained by training multilingual language models to map sentences from different languages into a shared semantic space. As such, they are subject to curse of multilinguality, a loss of monolingual representational accuracy due to parameter sharing. Another limitation of MSEs is the trade-off between different task performance: cross-lingual alignment training distorts the optimal monolingual structure of semantic spaces of individual languages, harming the utility of sentence embeddings in monolingual tasks; cross-lingual tasks, such as cross-lingual semantic similarity and zero-shot transfer for sentence classification, may also require conflicting cross-lingual alignment strategies. In this work, we address both issues by means of modular training of sentence encoders. We first train language-specific monolingual modules to mitigate negative interference between languages (i.e., the curse). We then align all non-English sentence embeddings to the English by training cross-lingual alignment adapters, preventing interference with monolingual specialization from the first step. We train the cross-lingual adapters with two different types of data to resolve the conflicting requirements of different cross-lingual tasks. Monolingual and cross-lingual results on semantic text similarity and relatedness, bitext mining and sentence classification show that our modular solution achieves better and more balanced performance across all the tasks compared to full-parameter training of monolithic multilingual sentence encoders, especially benefiting low-resource languages.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b3834dde-9ec6-4476-b441-41cb140e7accCited by top-tier papers2
- Unified Vision-Language Modeling via Concept Space AlignmentYifu Qiu, Paul-Ambroise Duquenne, Holger SchwenkICLR 2026
- Discriminating Form and Meaning in Multilingual Models with Minimal-Pair ABX TasksMaureen de Seyssel, Jie Chi, Skyler Seto, Maartje ter Hoeve et al.EMNLP 2025
Builds on16
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 2,496 citations
- Towards a Unified View of Parameter-Efficient Transfer LearningJunxian He, Chunting Zhou, Xuezhe Ma, Taylor Berg-Kirkpatrick et al.ICLR 2022 · 1,182 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- On Negative Interference in Multilingual Models: Findings and A Meta-Learning TreatmentZirui Wang, Zachary C. Lipton, Yulia TsvetkovEMNLP 2020 · 72 citations
Related papers
- Emu: Enhancing Multilingual Sentence Embeddings with Semantic SpecializationWataru Hirota, Yoshihiko Suhara, Behzad Golshan, Wang-Chiew TanAAAI 2020 · 5 citations
- Cross-lingual Sentence Embedding using Multi-Task LearningKoustava Goswami, Sourav Dutta, Haytham Assem, Theodorus Fransen et al.EMNLP 2021 · 9 citations
- Unsupervised Interlingual Semantic Representations from Sentence Embeddings for Zero-Shot Cross-Lingual TransferChanny Hong, Jaeyeon Lee, Jungkwon LeeAAAI 2020 · 1 citation
- Language-agnostic BERT Sentence EmbeddingFangxiaoyu Feng, Yinfei Yang, Daniel Cer, Naveen Arivazhagan et al.ACL 2022
- When Is Multilinguality a Curse? Language Modeling for 250 High- and Low-Resource LanguagesTyler A. Chang, Catherine Arnett, Zhuowen Tu, Ben BergenEMNLP 2024 · 12 citations
