Extreme Memorization via Scale of Initialization
Harsh Mehta, Ashok Cutkosky, Behnam Neyshabur
Abstract
We construct an experimental setup in which changing the scale of initialization strongly impacts the implicit regularization induced by SGD, interpolating from good generalization performance to completely memorizing the training set while making little progress on the test set. Moreover, we find that the extent and manner in which generalization ability is affected depends on the activation and loss function used, with sin activation demonstrating extreme memorization. In the case of the homogeneous ReLU activation, we show that this behavior can be attributed to the loss function. Our empirical investigation reveals that increasing the scale of initialization correlates with misalignment of representations and gradients across examples in the same class. This insight allows us to devise an alignment measure over gradients and representations which can capture this phenomenon. We demonstrate that our alignment measure correlates with generalization of deep models trained on image classification tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4d6d3f0c-5036-47e8-991a-7a1ffb4652a7Cited by top-tier papers5
- Gradient Descent on Two-layer Nets: Margin Maximization and Simplicity BiasKaifeng Lyu, Zhiyuan Li, Runzhe Wang, Sanjeev AroraNeurIPS 2021 · 94 citations
- Stochastic Training is Not Necessary for GeneralizationJonas Geiping, Micah Goldblum, Phillip Pope, Michael Moeller et al.ICLR 2022 · 83 citations
- Investigating Generalization by Controlling Normalized MarginAlexander R. Farhang, Jeremy D. Bernstein, Kushal Tirumala, Yang Liu et al.ICML 2022 · 6 citations
- Fast Training of Sinusoidal Neural Fields via Scaling InitializationTaesun Yeom, Sangyoon Lee, Jaeho LeeICLR 2025
- Over-Alignment vs Over-Fitting: The Role of Feature Learning Strength in GeneralizationTaesun Yeom, Taehyeok Ha, Jaeho LeeICML 2026
Builds on5
- Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional DomainsMatthew Tancik, Pratul P. Srinivasan, Ben Mildenhall, Sara Fridovich-Keil et al.NeurIPS 2020 · 4,036 citations
- Implicit Neural Representations with Periodic Activation FunctionsVincent Sitzmann, Julien N. P. Martel, Alexander W. Bergman, David B. Lindell et al.NeurIPS 2020 · 4,008 citations
- The Impact of Neural Network Overparameterization on Gradient Confusion and Stochastic Gradient DescentKarthik Abinav Sankararaman, Soham De, Zheng Xu, W. Ronny Huang et al.ICML 2020 · 122 citations
- Generalization bounds for deep convolutional neural networksPhilip M. Long, Hanie SedghiICLR 2020 · 102 citations
- Coherent Gradients: An Approach to Understanding Generalization in Gradient Descent-based OptimizationSatrajit ChatterjeeICLR 2020 · 60 citations
Related papers
- Catastrophic Fisher Explosion: Early Phase Fisher Matrix Impacts GeneralizationStanislaw Jastrzebski, Devansh Arpit, Oliver Åstrand, Giancarlo Kerg et al.ICML 2021 · 78 citations
- On the Origin of Implicit Regularization in Stochastic Gradient DescentSamuel L. Smith, Benoit Dherin, David G. T. Barrett, Soham DeICLR 2021 · 235 citations
- How Spurious Features are Memorized: Precise Analysis for Random and NTK FeaturesSimone Bombari, Marco MondelliICML 2024 · 10 citations
- Assessing Generalization of SGD via DisagreementYiding Jiang, Vaishnavh Nagarajan, Christina Baek, J. Zico KolterICLR 2022 · 134 citations
- On the Implicit Bias of Initialization Shape: Beyond Infinitesimal Mirror DescentShahar Azulay, Edward Moroshko, Mor Shpigel Nacson, Blake E. Woodworth et al.ICML 2021 · 85 citations
