mmBERT: A Modern Multilingual Encoder with Annealed Language Learning
Marc Marone, Orion Weller, William Fleshman, Eugene Yang, Dawn Lawrie, Benjamin Van Durme
Abstract
Encoder-only languages models are frequently used for a variety of standard machine learning tasks, including classification and retrieval. However, there has been a lack of recent research for encoder models, especially with respect to multilingual models. We introduce MMBERT, an encoder-only language model pretrained on 3T tokens of multilingual text in over 1800 languages. To build MMBERT we introduce several novel elements, including an inverse mask ratio schedule and an inverse temperature sampling ratio. We add over 1700 lowresource languages to the data mix only during the decay phase, showing that it boosts performance dramatically and maximizes the gains from the relatively small amount of training data. Despite only including these low-resource languages in the short decay phase we achieve similar classification performance to models like OpenAI's o3 and Google's Gemini 2.5 Pro. Overall, we show that MMBERT significantly outperforms the previous generation of models on classification and retrieval tasks -on both high and low-resource languages.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 29abed54-d1a0-45fe-b18f-a77ccc8d616fCited by top-tier papers4
- Seq vs Seq: An Open Suite of Paired Encoders and DecodersOrion Weller, Kathryn Ricci, Marc Marone, Antoine Chaffin et al.ICLR 2026 · 50 citations
- Uncovering the Latent Potential of Deep Intermediate RepresentationsArnesh Batra, Arush Gumber, Aniket Khandelwal, Jashn Khemani et al.ICML 2026 · 1 citation
- Learning to Route Languages for Multilingual Policy OptimizationGeyang Guo, Hiromi Wakaki, Yuki Mitsufuji, Alan Ritter et al.ICML 2026
- SwissGov-RSD: A Human-annotated, Cross-lingual Benchmark for Token-level Recognition of Semantic Differences Between Related DocumentsMichelle Wastl, Jannis Vamvas, Rico SennrichACL 2026
Builds on12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual GeneralisationJunjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig et al.ICML 2020 · 1,132 citations
- TIES-Merging: Resolving Interference When Merging ModelsPrateek Yadav, Derek Tam, Leshem Choshen, Colin A. Raffel et al.NeurIPS 2023 · 999 citations
- Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and InferenceBenjamin Warner, Antoine Chaffin, Benjamin Clavié, Orion Weller et al.ACL 2025 · 552 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
Related papers
- Evaluating the Cross-Lingual Effectiveness of Massively Multilingual Neural Machine TranslationAditya Siddhant, Melvin Johnson, Henry Tsai, Naveen Ari et al.AAAI 2020 · 74 citations
- BERTGen: Multi-task Generation through BERTFaidon Mitzalis, Ozan Caglayan, Pranava Madhyastha, Lucia SpeciaACL 2021
- From Zero to Hero: On the Limitations of Zero-Shot Language Transfer with Multilingual TransformersAnne Lauscher, Vinit Ravishankar, Ivan Vulic, Goran GlavasEMNLP 2020 · 235 citations
- Mixture of Languages: Improved Multilingual Encoders Through Language GroupingJoão Maria Janeiro, Belen Alastruey, Francisco Massa, Maha Elbayad et al.EMNLP 2025
- Zero-Shot Cross-Lingual Transfer of Neural Machine Translation with Multilingual Pretrained EncodersGuanhua Chen, Shuming Ma, Yun Chen, Li Dong et al.EMNLP 2021 · 30 citations
