The Curious Case of Benign Memorization
Sotiris Anagnostidis, Gregor Bachmann, Lorenzo Noci, Thomas Hofmann
Abstract
Despite the empirical advances of deep learning across a variety of learning tasks, our theoretical understanding of its success is still very restricted. One of the key challenges is the overparametrized nature of modern models, enabling complete overfitting of the data even if the labels are randomized, i.e. networks can completely memorize all given patterns. While such a memorization capacity seems worrisome, in this work we show that under training protocols that include data augmentation, neural networks learn to memorize entirely random labels in a benign way, i.e. they learn embeddings that lead to highly non-trivial performance under nearest neighbour probing. We demonstrate that deep models have the surprising ability to separate noise from signal by distributing the task of memorization and feature learning to different layers. As a result, only the very last layers are used for memorization, while preceding layers encode performant features which remain largely unaffected by the label noise. We explore the intricate role of the augmentations used for training and identify a memorization-generalization trade-off in terms of their diversity, marking a clear distinction to all previous works. Finally, we give a first explanation for the emergence of benign memorization by showing that malign memorization under data augmentation is infeasible due to the insufficient capacity of the model for the increased sample size. As a consequence, the network is forced to leverage the correlated nature of the augmentations and as a result learns meaningful features. To complete the picture, a better theory of feature learning in deep neural networks is required to fully understand the origins of this phenomenon.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e088d1b0-83cd-4f34-894d-596db3c11a78Cited by top-tier papers6
- Scaling MLPs: A Tale of Inductive BiasGregor Bachmann, Sotiris Anagnostidis, Thomas HofmannNeurIPS 2023 · 71 citations
- Random Teachers are Good TeachersFelix Sarnthein, Gregor Bachmann, Sotiris Anagnostidis, Thomas HofmannICML 2023 · 8 citations
- Spatio-Temporal Crop Aggregation for Video Representation LearningSepehr Sameni, Simon Jenni, Paolo FavaroICCV 2023 · 4 citations
- Causal Estimation of Memorisation ProfilesPietro Lesci, Clara Meister, Thomas Hofmann, Andreas Vlachos et al.ACL 2024 · 3 citations
- ELDET: Early-Learning Distillation with Noisy Labels for Object DetectionDongmin Choi, Sangbin Lee, EungGu Yun, Jonghyuk Baek et al.NeurIPS 2025 · 1 citation
Builds on28
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 4,453 citations
- Barlow Twins: Self-Supervised Learning via Redundancy ReductionJure Zbontar, Li Jing, Ishan Misra, Yann LeCun et al.ICML 2021 · 2,942 citations
Related papers
- Random Label Prediction Heads for Studying Memorization in Deep Neural NetworksMarlon Becker, Jonas Konrad, Luis Garcia Rodriguez, Benjamin RisseICLR 2026
- How does the Memorization of Neural Networks Impact Adversarial Robust Models?Han Xu, Xiaorui Liu, Wentao Wang, Zitao Liu et al.KDD 2023 · 1 citation
- Exploring Memorization in Adversarial TrainingYinpeng Dong, Ke Xu, Xiao Yang, Tianyu Pang et al.ICLR 2022 · 84 citations
- A Law of Data Reconstruction for Random Features (And Beyond)Leonardo Iurada, Simone Bombari, Tatiana Tommasi, Marco MondelliICLR 2026 · 3 citations
- How Spurious Features are Memorized: Precise Analysis for Random and NTK FeaturesSimone Bombari, Marco MondelliICML 2024 · 10 citations
