Rare Words: A Major Problem for Contextualized Embeddings and How to Fix it by Attentive Mimicking
Timo Schick, Hinrich Schütze
摘要
Pretraining deep neural network architectures with a language modeling objective has brought large improvements for many natural language processing tasks. Exemplified by BERT, a recently proposed such architecture, we demonstrate that despite being trained on huge amounts of data, deep language models still struggle to understand rare words. To fix this problem, we adapt Attentive Mimicking, a method that was designed to explicitly learn embeddings for rare words, to deep language models. In order to make this possible, we introduce one-token approximation, a procedure that enables us to use Attentive Mimicking even when the underlying language model uses subword-based tokenization, i.e., it does not assign embeddings to all words. To evaluate our method, we create a novel dataset that tests the ability of language models to capture semantic properties of words without any task-specific fine-tuning. Using this dataset, we show that adding our adapted version of Attentive Mimicking to BERT does substantially improve its understanding of rare words.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Few-Shot Text Generation with Natural Language InstructionsTimo Schick, Hinrich SchützeEMNLP 2021 · 被引用 101 次
- Towards Open-World Feature Extrapolation: An Inductive Graph Learning ApproachQitian Wu, Chenxiao Yang, Junchi YanNeurIPS 2021 · 被引用 39 次
- Improving Tokenisation by Alternative Treatment of SpacesEdward Gow-Smith, Harish Tayyar Madabushi, Carolina Scarton, Aline VillavicencioEMNLP 2022 · 被引用 6 次
- Graph-based Relation Mining for Context-free Out-of-vocabulary Word Embedding LearningZiran Liang, Yuyin Lu, Hegang Chen, Yanghui RaoACL 2023 · 被引用 5 次
- WSPAlign: Word Alignment Pre-training via Large-Scale Weakly Supervised Span PredictionQiyu Wu, Masaaki Nagata, Yoshimasa TsuruokaACL 2023 · 被引用 3 次
相关 Paper
- BERTRAM: Improved Word Embeddings Have Big Impact on Contextualized Model PerformanceTimo Schick, Hinrich SchützeACL 2020 · 被引用 1 次
- Imputing Out-of-Vocabulary Embeddings with LOVE Makes LanguageModels Robust with Little CostLihu Chen, Gaël Varoquaux, Fabian M. SuchanekACL 2022
- Seeing Clearly, Reasoning Confidently: Plug-and-Play Remedies for Vision Language Model BlindnessXin Hu, Haomiao Ni, Yunbei Zhang, Jihun Hamm 等CVPR 2026 · 被引用 1 次
- NextLevelBERT: Masked Language Modeling with Higher-Level Representations for Long DocumentsTamara Czinczoll, Christoph Hönes, Maximilian Schall, Gerard de MeloACL 2024 · 被引用 4 次
- Memorisation versus Generalisation in Pre-trained Language ModelsMichael Tänzer, Sebastian Ruder, Marek ReiACL 2022 · 被引用 59 次
