The Curious Case of Benign Memorization
Sotiris Anagnostidis, Gregor Bachmann, Lorenzo Noci, Thomas Hofmann
摘要
Despite the empirical advances of deep learning across a variety of learning tasks, our theoretical understanding of its success is still very restricted. One of the key challenges is the overparametrized nature of modern models, enabling complete overfitting of the data even if the labels are randomized, i.e. networks can completely memorize all given patterns. While such a memorization capacity seems worrisome, in this work we show that under training protocols that include data augmentation, neural networks learn to memorize entirely random labels in a benign way, i.e. they learn embeddings that lead to highly non-trivial performance under nearest neighbour probing. We demonstrate that deep models have the surprising ability to separate noise from signal by distributing the task of memorization and feature learning to different layers. As a result, only the very last layers are used for memorization, while preceding layers encode performant features which remain largely unaffected by the label noise. We explore the intricate role of the augmentations used for training and identify a memorization-generalization trade-off in terms of their diversity, marking a clear distinction to all previous works. Finally, we give a first explanation for the emergence of benign memorization by showing that malign memorization under data augmentation is infeasible due to the insufficient capacity of the model for the increased sample size. As a consequence, the network is forced to leverage the correlated nature of the augmentations and as a result learns meaningful features. To complete the picture, a better theory of feature learning in deep neural networks is required to fully understand the origins of this phenomenon.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Scaling MLPs: A Tale of Inductive BiasGregor Bachmann, Sotiris Anagnostidis, Thomas HofmannNeurIPS 2023 · 被引用 71 次
- Random Teachers are Good TeachersFelix Sarnthein, Gregor Bachmann, Sotiris Anagnostidis, Thomas HofmannICML 2023 · 被引用 8 次
- Spatio-Temporal Crop Aggregation for Video Representation LearningSepehr Sameni, Simon Jenni, Paolo FavaroICCV 2023 · 被引用 4 次
- Causal Estimation of Memorisation ProfilesPietro Lesci, Clara Meister, Thomas Hofmann, Andreas Vlachos 等ACL 2024 · 被引用 3 次
- ELDET: Early-Learning Distillation with Noisy Labels for Object DetectionDongmin Choi, Sangbin Lee, EungGu Yun, Jonghyuk Baek 等NeurIPS 2025 · 被引用 1 次
它引用的顶会 Paper28
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 被引用 4,453 次
- Barlow Twins: Self-Supervised Learning via Redundancy ReductionJure Zbontar, Li Jing, Ishan Misra, Yann LeCun 等ICML 2021 · 被引用 2,942 次
相关 Paper
- Random Label Prediction Heads for Studying Memorization in Deep Neural NetworksMarlon Becker, Jonas Konrad, Luis Garcia Rodriguez, Benjamin RisseICLR 2026
- How does the Memorization of Neural Networks Impact Adversarial Robust Models?Han Xu, Xiaorui Liu, Wentao Wang, Zitao Liu 等KDD 2023 · 被引用 1 次
- Exploring Memorization in Adversarial TrainingYinpeng Dong, Ke Xu, Xiao Yang, Tianyu Pang 等ICLR 2022 · 被引用 84 次
- A Law of Data Reconstruction for Random Features (And Beyond)Leonardo Iurada, Simone Bombari, Tatiana Tommasi, Marco MondelliICLR 2026 · 被引用 3 次
- How Spurious Features are Memorized: Precise Analysis for Random and NTK FeaturesSimone Bombari, Marco MondelliICML 2024 · 被引用 10 次
