MEXMA: Token-level objectives improve sentence representations
João Maria Janeiro, Benjamin Piwowarski, Patrick Gallinari, Loïc Barrault
Abstract
Current pre-trained cross-lingual sentence encoders approaches use sentence-level objectives only. This can lead to loss of information, especially for tokens, which then degrades the sentence representation. We propose MEXMA, a novel approach that integrates both sentence-level and token-level objectives. The sentence representation in one language is used to predict masked tokens in another language, with both the sentence representation and all tokens directly updating the encoder. We show that adding token-level objectives greatly improves the sentence representation quality across several tasks. Our approach outperforms current pre-trained cross-lingual sentence encoders on bi-text mining as well as several downstream tasks. We also analyse the information encoded in our tokens, and how the sentence representation is built from them.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 06ddece9-c73e-4cfb-85ce-2b6df6cb90b5Cited by top-tier papers2
- Middle-Layer Representation Alignment for Cross-Lingual Transfer in Fine-Tuned LLMsDanni Liu, Jan NiehuesACL 2025 · 23 citations
- Mixture of Languages: Improved Multilingual Encoders Through Language GroupingJoão Maria Janeiro, Belen Alastruey, Francisco Massa, Maha Elbayad et al.EMNLP 2025
Builds on11
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
- XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual GeneralisationJunjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig et al.ICML 2020 · 1,132 citations
- Vision Transformers Need RegistersTimothée Darcet, Maxime Oquab, Julien Mairal, Piotr BojanowskiICLR 2024 · 769 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
Related papers
- Modeling Sequential Sentence Relation to Improve Cross-lingual Dense RetrievalShunyu Zhang, Yaobo Liang, Ming Gong, Daxin Jiang et al.ICLR 2023 · 1 citation
- Dual-Alignment Pre-training for Cross-lingual Sentence EmbeddingZiheng Li, Shaohan Huang, Zihan Zhang, Zhi-Hong Deng et al.ACL 2023 · 5 citations
- Representation Deficiency in Masked Language ModelingYu Meng, Jitin Krishnan, Sinong Wang, Qifan Wang et al.ICLR 2024 · 12 citations
- UniLMv2: Pseudo-Masked Language Models for Unified Language Model Pre-TrainingHangbo Bao, Li Dong, Furu Wei, Wenhui Wang et al.ICML 2020 · 423 citations
- RetroMAE-2: Duplex Masked Auto-Encoder For Pre-Training Retrieval-Oriented Language ModelsZheng Liu, Shitao Xiao, Yingxia Shao, Zhao CaoACL 2023 · 7 citations
