FOCUS: Effective Embedding Initialization for Monolingual Specialization of Multilingual Models
Konstantin Dobler, Gerard de Melo
Abstract
Using model weights pretrained on a high-resource language as a warm start can reduce the need for data and compute to obtain high-quality language models for other, especially low-resource, languages. However, if we want to use a new tokenizer specialized for the target language, we cannot transfer the source model’s embedding matrix. In this paper, we propose FOCUS - Fast Overlapping Token Combinations Using Sparsemax, a novel embedding initialization method that effectively initializes the embedding matrix for a new tokenizer based on information in the source model’s embedding matrix. FOCUS represents newly added tokens as combinations of tokens in the overlap of the source and target vocabularies. The overlapping tokens are selected based on semantic similarity in an auxiliary static token embedding space. We focus our study on using the multilingual XLM-R as a source model and empirically show that FOCUS outperforms random initialization and previous work on language modeling and on a range of downstream tasks (NLI, QA, and NER). We publish our model checkpoints and code on GitHub.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 43649475-4137-4b64-8556-0be194be58bcCited by top-tier papers17
- Universal Cross-Tokenizer Distillation via Approximate Likelihood MatchingBenjamin Minixhofer, Ivan Vulic, Edoardo Maria PontiNeurIPS 2025 · 48 citations
- Zero-Shot Tokenizer TransferBenjamin Minixhofer, Edoardo Maria Ponti, Ivan VulicNeurIPS 2024 · 37 citations
- One Tokenizer To Rule Them All: Emergent Language Plasticity via Multilingual TokenizersDiana Abagyan, Alejandro Salamanca, Andrés Felipe Cruz-Salinas, Kris Cao et al.ACL 2026 · 11 citations
- Sampling from Your Language Model One Byte at a TimeJonathan Hayase, Alisa Liu, Noah Smith, Sewoong OhICML 2026 · 9 citations
- Token Distillation: Attention-Aware Input Embeddings for New TokensKonstantin Dobler, Desmond Elliott, Gerard de MeloICLR 2026 · 7 citations
Builds on8
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- Expanding Pretrained Models to Thousands More Languages via Lexicon-based AdaptationXinyi Wang, Sebastian Ruder, Graham NeubigACL 2022 · 73 citations
- On the Cross-lingual Transferability of Monolingual RepresentationsMikel Artetxe, Sebastian Ruder, Dani YogatamaACL 2020 · 57 citations
- MAD-X: An Adapter-Based Framework for Multi-Task Cross-Lingual TransferJonas Pfeiffer, Ivan Vulic, Iryna Gurevych, Sebastian RuderEMNLP 2020 · 36 citations
- ABSent: Cross-Lingual Sentence Representation Mapping with Bidirectional GANsZuohui Fu, Yikun Xian, Shijie Geng, Yingqiang Ge et al.AAAI 2020 · 20 citations
Related papers
- Overlap-based Vocabulary Generation Improves Cross-lingual Transfer Among Related LanguagesVaidehi Patil, Partha P. Talukdar, Sunita SarawagiACL 2022 · 39 citations
- Beyond Offline Mapping: Learning Cross-lingual Word Embeddings through Context AnchoringAitor Ormazabal, Mikel Artetxe, Aitor Soroa, Gorka Labaka et al.ACL 2021
- TEMA: Token Embeddings Mapping for Enriching Low-Resource Language ModelsRodolfo Zevallos, Núria Bel, Mireia FarrúsEMNLP 2024
- Revisiting Tri-training of Dependency ParsersJoachim Wagner, Jennifer FosterEMNLP 2021
- UNKs Everywhere: Adapting Multilingual Language Models to New ScriptsJonas Pfeiffer, Ivan Vulic, Iryna Gurevych, Sebastian RuderEMNLP 2021 · 3 citations
