Rare Words: A Major Problem for Contextualized Embeddings and How to Fix it by Attentive Mimicking
Timo Schick, Hinrich Schütze
Abstract
Pretraining deep neural network architectures with a language modeling objective has brought large improvements for many natural language processing tasks. Exemplified by BERT, a recently proposed such architecture, we demonstrate that despite being trained on huge amounts of data, deep language models still struggle to understand rare words. To fix this problem, we adapt Attentive Mimicking, a method that was designed to explicitly learn embeddings for rare words, to deep language models. In order to make this possible, we introduce one-token approximation, a procedure that enables us to use Attentive Mimicking even when the underlying language model uses subword-based tokenization, i.e., it does not assign embeddings to all words. To evaluate our method, we create a novel dataset that tests the ability of language models to capture semantic properties of words without any task-specific fine-tuning. Using this dataset, we show that adding our adapted version of Attentive Mimicking to BERT does substantially improve its understanding of rare words.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- Few-Shot Text Generation with Natural Language InstructionsTimo Schick, Hinrich SchützeEMNLP 2021 · 101 citations
- Towards Open-World Feature Extrapolation: An Inductive Graph Learning ApproachQitian Wu, Chenxiao Yang, Junchi YanNeurIPS 2021 · 39 citations
- Improving Tokenisation by Alternative Treatment of SpacesEdward Gow-Smith, Harish Tayyar Madabushi, Carolina Scarton, Aline VillavicencioEMNLP 2022 · 6 citations
- Graph-based Relation Mining for Context-free Out-of-vocabulary Word Embedding LearningZiran Liang, Yuyin Lu, Hegang Chen, Yanghui RaoACL 2023 · 5 citations
- WSPAlign: Word Alignment Pre-training via Large-Scale Weakly Supervised Span PredictionQiyu Wu, Masaaki Nagata, Yoshimasa TsuruokaACL 2023 · 3 citations
Related papers
- BERTRAM: Improved Word Embeddings Have Big Impact on Contextualized Model PerformanceTimo Schick, Hinrich SchützeACL 2020 · 1 citation
- Imputing Out-of-Vocabulary Embeddings with LOVE Makes LanguageModels Robust with Little CostLihu Chen, Gaël Varoquaux, Fabian M. SuchanekACL 2022
- Seeing Clearly, Reasoning Confidently: Plug-and-Play Remedies for Vision Language Model BlindnessXin Hu, Haomiao Ni, Yunbei Zhang, Jihun Hamm et al.CVPR 2026 · 1 citation
- NextLevelBERT: Masked Language Modeling with Higher-Level Representations for Long DocumentsTamara Czinczoll, Christoph Hönes, Maximilian Schall, Gerard de MeloACL 2024 · 4 citations
- Memorisation versus Generalisation in Pre-trained Language ModelsMichael Tänzer, Sebastian Ruder, Marek ReiACL 2022 · 59 citations
