On the geometry of generalization and memorization in deep neural networks
Cory Stephenson, Suchismita Padhy, Abhinav Ganesh, Yue Hui, Hanlin Tang, SueYeon Chung
Abstract
Understanding how large neural networks avoid memorizing training data is key to explaining their high generalization performance. To examine the structure of when and where memorization occurs in a deep network, we use a recently developed replica-based mean field theoretic geometric analysis method. We find that all layers preferentially learn from examples which share features, and link this behavior to generalization performance. Memorization predominately occurs in the deeper layers, due to decreasing object manifolds' radius and dimension, whereas early layers are minimally affected. This predicts that generalization can be restored by reverting the final few layer weights to earlier epochs before significant memorization occurred, which is confirmed by the experiments. Additionally, by studying generalization under different model sizes, we reveal the connection between the double descent phenomenon and the underlying model geometry. Finally, analytical analysis shows that networks avoid memorization early in training because close to initialization, the gradient contribution from permuted examples are small. These findings provide quantitative evidence for the structure of memorization across layers of a deep neural network, the drivers for such structure, and its connection to manifold geometric properties.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8c0aaf0d-84c9-4376-af01-b43bbee30484Cited by top-tier papers31
- Deduplicating Training Data Makes Language Models BetterKatherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang et al.ACL 2022 · 844 citations
- Fast Machine Unlearning without Retraining through Selective Synaptic DampeningJack Foster, Stefan Schoepf, Alexandra BrintrupAAAI 2024 · 208 citations
- Deep Learning Through the Lens of Example DifficultyRobert J. N. Baldock, Hartmut Maennel, Behnam NeyshaburNeurIPS 2021 · 204 citations
- Balancing Discriminability and Transferability for Source-Free Domain AdaptationJogendra Nath Kundu, Akshay R. Kulkarni, Suvaansh Bhambri, Deepesh Mehta et al.ICML 2022 · 110 citations
- On Memorization in Probabilistic Deep Generative ModelsGerrit J. J. van den Burg, Christopher K. I. WilliamsNeurIPS 2021 · 92 citations
Builds on3
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang et al.ICLR 2020 · 1,108 citations
- Coherent Gradients: An Approach to Understanding Generalization in Gradient Descent-based OptimizationSatrajit ChatterjeeICLR 2020 · 60 citations
- Emergence of Separable Manifolds in Deep Language RepresentationsJonathan Mamou, Hang Le, Miguel Del Rio, Cory Stephenson et al.ICML 2020 · 52 citations
Related papers
- Model, sample, and epoch-wise descents: exact solution of gradient flow in the random feature modelAntoine Bodin, Nicolas MacrisNeurIPS 2021 · 19 citations
- Multi-scale Feature Learning Dynamics: Insights for Double DescentMohammad Pezeshki, Amartya Mitra, Yoshua Bengio, Guillaume LajoieICML 2022 · 33 citations
- Learning Curves for Deep Structured Gaussian Feature ModelsJacob A. Zavatone-Veth, Cengiz PehlevanNeurIPS 2023 · 15 citations
- Can Neural Nets Learn the Same Model Twice? Investigating Reproducibility and Double Descent from the Decision Boundary PerspectiveGowthami Somepalli, Liam Fowl, Arpit Bansal, Ping-Yeh Chiang et al.CVPR 2022
- A Geometric Framework for Understanding Memorization in Generative ModelsBrendan Leigh Ross, Hamidreza Kamkari, Tongzi Wu, Rasa Hosseinzadeh et al.ICLR 2025
