mmBERT: A Modern Multilingual Encoder with Annealed Language Learning
Marc Marone, Orion Weller, William Fleshman, Eugene Yang, Dawn Lawrie, Benjamin Van Durme
摘要
Encoder-only languages models are frequently used for a variety of standard machine learning tasks, including classification and retrieval. However, there has been a lack of recent research for encoder models, especially with respect to multilingual models. We introduce MMBERT, an encoder-only language model pretrained on 3T tokens of multilingual text in over 1800 languages. To build MMBERT we introduce several novel elements, including an inverse mask ratio schedule and an inverse temperature sampling ratio. We add over 1700 lowresource languages to the data mix only during the decay phase, showing that it boosts performance dramatically and maximizes the gains from the relatively small amount of training data. Despite only including these low-resource languages in the short decay phase we achieve similar classification performance to models like OpenAI's o3 and Google's Gemini 2.5 Pro. Overall, we show that MMBERT significantly outperforms the previous generation of models on classification and retrieval tasks -on both high and low-resource languages.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Seq vs Seq: An Open Suite of Paired Encoders and DecodersOrion Weller, Kathryn Ricci, Marc Marone, Antoine Chaffin 等ICLR 2026 · 被引用 50 次
- Uncovering the Latent Potential of Deep Intermediate RepresentationsArnesh Batra, Arush Gumber, Aniket Khandelwal, Jashn Khemani 等ICML 2026 · 被引用 1 次
- Learning to Route Languages for Multilingual Policy OptimizationGeyang Guo, Hiromi Wakaki, Yuki Mitsufuji, Alan Ritter 等ICML 2026
- SwissGov-RSD: A Human-annotated, Cross-lingual Benchmark for Token-level Recognition of Semantic Differences Between Related DocumentsMichelle Wastl, Jannis Vamvas, Rico SennrichACL 2026
它引用的顶会 Paper12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual GeneralisationJunjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig 等ICML 2020 · 被引用 1,132 次
- TIES-Merging: Resolving Interference When Merging ModelsPrateek Yadav, Derek Tam, Leshem Choshen, Colin A. Raffel 等NeurIPS 2023 · 被引用 999 次
- Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and InferenceBenjamin Warner, Antoine Chaffin, Benjamin Clavié, Orion Weller 等ACL 2025 · 被引用 552 次
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
相关 Paper
- Evaluating the Cross-Lingual Effectiveness of Massively Multilingual Neural Machine TranslationAditya Siddhant, Melvin Johnson, Henry Tsai, Naveen Ari 等AAAI 2020 · 被引用 74 次
- BERTGen: Multi-task Generation through BERTFaidon Mitzalis, Ozan Caglayan, Pranava Madhyastha, Lucia SpeciaACL 2021
- From Zero to Hero: On the Limitations of Zero-Shot Language Transfer with Multilingual TransformersAnne Lauscher, Vinit Ravishankar, Ivan Vulic, Goran GlavasEMNLP 2020 · 被引用 235 次
- Mixture of Languages: Improved Multilingual Encoders Through Language GroupingJoão Maria Janeiro, Belen Alastruey, Francisco Massa, Maha Elbayad 等EMNLP 2025
- Zero-Shot Cross-Lingual Transfer of Neural Machine Translation with Multilingual Pretrained EncodersGuanhua Chen, Shuming Ma, Yun Chen, Li Dong 等EMNLP 2021 · 被引用 30 次
