Improving Language Plasticity via Pretraining with Active Forgetting
Yihong Chen, Kelly Marchisio, Roberta Raileanu, David Ifeoluwa Adelani, Pontus Lars Erik Saito Stenetorp, Sebastian Riedel, Mikel Artetxe
Abstract
Pretrained language models (PLMs) are today the primary model for natural language processing. Despite their impressive downstream performance, it can be difficult to apply PLMs to new languages, a barrier to making their capabilities universally accessible. While prior work has shown it possible to address this issue by learning a new embedding layer for the new language, doing so is both data and compute inefficient. We propose to use an active forgetting mechanism during pretraining, as a simple way of creating PLMs that can quickly adapt to new languages. Concretely, by resetting the embedding layer every K updates during pretraining, we encourage the PLM to improve its ability of learning new embeddings within limited number of updates, similar to a meta-learning effect. Experiments with RoBERTa show that models pretrained with our forgetting mechanism not only demonstrate faster convergence during language adaptation, but also outperform standard ones in a low-data regime, particularly for languages that are distant from English. Code will be available at https://github.com/ facebookresearch/language-model-plasticity .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9a4e5128-f1d5-4b87-9b89-44bbd7fbaf87Cited by top-tier papers20
- Zero-Shot Tokenizer TransferBenjamin Minixhofer, Edoardo Maria Ponti, Ivan VulicNeurIPS 2024 · 37 citations
- Breaking Physical and Linguistic Borders: Multilingual Federated Prompt Tuning for Low-Resource LanguagesWanru Zhao, Yihong Chen, Royson Lee, Xinchi Qiu et al.ICLR 2024 · 21 citations
- Scaling Embedding Layers in Language ModelsDa Yu, Edith Cohen, Badih Ghazi, Yangsibo Huang et al.NeurIPS 2025 · 21 citations
- Co-occurrence is not Factual Association in Language ModelsXiao Zhang, Miao Li, Ji WuNeurIPS 2024 · 15 citations
- Understanding Language Prior of LVLMs by Contrasting Chain-of-EmbeddingLin Long, Changdae Oh, Seongheon Park, Sharon LiICLR 2026 · 14 citations
Builds on24
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- Fast Model Editing at ScaleEric Mitchell, Charles Lin, Antoine Bosselut, Chelsea Finn et al.ICLR 2022 · 527 citations
Related papers
- One Tokenizer To Rule Them All: Emergent Language Plasticity via Multilingual TokenizersDiana Abagyan, Alejandro Salamanca, Andrés Felipe Cruz-Salinas, Kris Cao et al.ACL 2026 · 11 citations
- On the Effectiveness of Adapter-based Tuning for Pretrained Language Model AdaptationRuidan He, Linlin Liu, Hai Ye, Qingyu Tan et al.ACL 2021
- On the Usage of Continual Learning for Out-of-Distribution Generalization in Pre-trained Language Models of CodeMartin Weyssow, Xin Zhou, Kisub Kim, David Lo et al.FSE 2023 · 9 citations
- Analyzing and Reducing the Performance Gap in Cross-Lingual Transfer with Fine-tuning Slow and FastYiduo Guo, Yaobo Liang, Dongyan Zhao, Bing Liu et al.ACL 2023
- Emergent Abilities of Large Language Models under Continued Pre-training for Language AdaptationAhmed Elhady, Eneko Agirre, Mikel ArtetxeACL 2025
