Memorisation versus Generalisation in Pre-trained Language Models
Michael Tänzer, Sebastian Ruder, Marek Rei
Abstract
State-of-the-art pre-trained language models have been shown to memorise facts and perform well with limited amounts of training data. To gain a better understanding of how these models learn, we study their generalisation and memorisation capabilities in noisy and low-resource scenarios. We find that the training of these models is almost unaffected by label noise and that it is possible to reach near-optimal results even on extremely noisy datasets. However, our experiments also show that they mainly learn from high-frequency patterns and largely fail when tested on low-resource tasks such as few-shot learning and rare entity recognition. To mitigate such limitations, we propose an extension based on prototypical networks that improves performance in low-resource named entity recognition tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers15
- How Do Large Language Models Acquire Factual Knowledge During Pretraining?Hoyeon Chang, Jinho Park, Seonghyeon Ye, Sohee Yang et al.NeurIPS 2024 · 124 citations
- DyVal: Dynamic Evaluation of Large Language Models for Reasoning TasksKaijie Zhu, Jiaao Chen, Jindong Wang, Neil Zhenqiang Gong et al.ICLR 2024 · 92 citations
- Decoupling Knowledge from Memorization: Retrieval-augmented Prompt LearningXiang Chen, Lei Li, Ningyu Zhang, Xiaozhuan Liang et al.NeurIPS 2022 · 68 citations
- Moderate-fitting as a Natural Backdoor Defender for Pre-trained Language ModelsBiru Zhu, Yujia Qin, Ganqu Cui, Yangyi Chen et al.NeurIPS 2022 · 29 citations
- Weaker Than You Think: A Critical Look at Weakly Supervised LearningDawei Zhu, Xiaoyu Shen, Marius Mosbach, Andreas Stephan et al.ACL 2023 · 13 citations
Builds on1
Related papers
- Few-Shot Named Entity Recognition: An Empirical Baseline StudyJiaxin Huang, Chunyuan Li, Krishan Subudhi, Damien Jose et al.EMNLP 2021 · 97 citations
- Adversity-aware Few-shot Named Entity Recognition via Augmentation LearningLi Huang, Haowen Liu, Qiang Gao, Jiajing Yu et al.AAAI 2025 · 1 citation
- Empirical Analysis of Unlabeled Entity Problem in Named Entity RecognitionYangming Li, Lemao Liu, Shuming ShiICLR 2021 · 72 citations
- Prototypical Fine-Tuning: Towards Robust Performance under Varying Data SizesYiqiao Jin, Xiting Wang, Yaru Hao, Yizhou Sun et al.AAAI 2023 · 15 citations
- Improving Low-Resource Languages in Pre-Trained Multilingual Language ModelsViktor Hangya, Hossain Shaikh Saadi, Alexander FraserEMNLP 2022 · 17 citations
