BERTRAM: Improved Word Embeddings Have Big Impact on Contextualized Model Performance
Timo Schick, Hinrich Schütze
Abstract
Pretraining deep language models has led to large performance gains in NLP. Despite this success, Schick and Schütze (2020) recently showed that these models struggle to understand rare words. For static word embeddings, this problem has been addressed by separately learning representations for rare words. In this work, we transfer this idea to pretrained language models: We introduce BERTRAM, a powerful architecture based on BERT that is capable of inferring high-quality embeddings for rare words that are suitable as input representations for deep language models. This is achieved by enabling the surface form and contexts of a word to interact with each other in a deep architecture. Integrating BERTRAM into BERT leads to large performance increases due to improved representations of rare and medium frequency words on both a rare word probing task and three downstream tasks. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a98cf196-54f0-4ced-9c12-1de340167815Cited by top-tier papers6
- Zero-Shot Tokenizer TransferBenjamin Minixhofer, Edoardo Maria Ponti, Ivan VulicNeurIPS 2024 · 37 citations
- Token Distillation: Attention-Aware Input Embeddings for New TokensKonstantin Dobler, Desmond Elliott, Gerard de MeloICLR 2026 · 7 citations
- READIN: A Chinese Multi-Task Benchmark with Realistic and Diverse Input NoisesChenglei Si, Zhengyan Zhang, Yingfa Chen, Xiaozhi Wang et al.ACL 2023 · 1 citation
- Grounded Compositional Outputs for Adaptive Language ModelingNikolaos Pappas, Phoebe Mulcaire, Noah A. SmithEMNLP 2020 · 1 citation
- Retrofitting Large Language Models with Dynamic TokenizationDarius Feher, Ivan Vulic, Benjamin MinixhoferACL 2025
Related papers
- Rare Words: A Major Problem for Contextualized Embeddings and How to Fix it by Attentive MimickingTimo Schick, Hinrich SchützeAAAI 2020 · 106 citations
- Superbizarre Is Not Superb: Derivational Morphology Improves BERT's Interpretation of Complex WordsValentin Hofmann, Janet B. Pierrehumbert, Hinrich SchützeACL 2021
- Better Language Model with Hypernym Class PredictionHe Bai, Tong Wang, Alessandro Sordoni, Peng ShiACL 2022
- Imputing Out-of-Vocabulary Embeddings with LOVE Makes LanguageModels Robust with Little CostLihu Chen, Gaël Varoquaux, Fabian M. SuchanekACL 2022
- Inducing Relational Knowledge from BERTZied Bouraoui, José Camacho-Collados, Steven SchockaertAAAI 2020 · 183 citations
