An Empirical Study of Memorization in NLP
Xiaosen Zheng, Jing Jiang
摘要
A recent study by Feldman (2020) proposed a long-tail theory to explain the memorization behavior of deep learning models. However, memorization has not been empirically verified in the context of NLP, a gap addressed by this work. In this paper, we use three different NLP tasks to check if the long-tail theory holds. Our experiments demonstrate that top-ranked memorized training instances are likely atypical, and removing the top-memorized training instances leads to a more serious drop in test accuracy compared with removing training instances randomly. Furthermore, we develop an attribution method to better understand why a training instance is memorized. We empirically show that our memorization attribution method is faithful, and share our interesting finding that the top-memorized parts of a training instance tend to be features negatively correlated with the class label.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Decoupling Knowledge from Memorization: Retrieval-augmented Prompt LearningXiang Chen, Lei Li, Ningyu Zhang, Xiaozhuan Liang 等NeurIPS 2022 · 被引用 68 次
- Intriguing Properties of Data Attribution on Diffusion ModelsXiaosen Zheng, Tianyu Pang, Chao Du, Jing Jiang 等ICLR 2024 · 被引用 41 次
- Causal Estimation of Memorisation ProfilesPietro Lesci, Clara Meister, Thomas Hofmann, Andreas Vlachos 等ACL 2024 · 被引用 3 次
- Mitigating Frequency Bias and Anisotropy in Language Model Pre-Training with Syntactic SmoothingRichard Diehl Martinez, Zébulon Goriely, Andrew Caines, Paula Buttery 等EMNLP 2024 · 被引用 1 次
- Memorisation Cartography: Mapping out the Memorisation-Generalisation Continuum in Neural Machine TranslationVerna Dankers, Ivan Titov, Dieuwke HupkesEMNLP 2023
它引用的顶会 Paper9
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 被引用 5,137 次
- Generalization through Memorization: Nearest Neighbor Language ModelsUrvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer 等ICLR 2020 · 被引用 1,038 次
- Estimating Training Data Influence by Tracing Gradient DescentGarima Pruthi, Frederick Liu, Satyen Kale, Mukund SundararajanNeurIPS 2020 · 被引用 784 次
- What Neural Networks Memorize and Why: Discovering the Long Tail via Influence EstimationVitaly Feldman, Chiyuan ZhangNeurIPS 2020 · 被引用 674 次
- Nearest Neighbor Machine TranslationUrvashi Khandelwal, Angela Fan, Dan Jurafsky, Luke Zettlemoyer 等ICLR 2021 · 被引用 323 次
相关 Paper
- Memorizing Long-tail Data Can Help Generalization Through CompositionMo Zhou, Haoyang Ma, Rong GeICLR 2026
- Counterfactual Memorization in Neural Language ModelsChiyuan Zhang, Daphne Ippolito, Katherine Lee, Matthew Jagielski 等NeurIPS 2023 · 被引用 184 次
- Memorization Through the Lens of Curvature of Loss Function Around SamplesIsha Garg, Deepak Ravikumar, Kaushik RoyICML 2024 · 被引用 25 次
- Meta-LMTC: Meta-Learning for Large-Scale Multi-Label Text ClassificationRan Wang, Xi'ao Su, Siyu Long, Xinyu Dai 等EMNLP 2021 · 被引用 10 次
- How Do Large Language Models Acquire Factual Knowledge During Pretraining?Hoyeon Chang, Jinho Park, Seonghyeon Ye, Sohee Yang 等NeurIPS 2024 · 被引用 124 次
