Over-Alignment vs Over-Fitting: The Role of Feature Learning Strength in Generalization
Taesun Yeom, Taehyeok Ha, Jaeho Lee
Abstract
Feature learning strength (FLS), i.e., the inverse of the effective output scaling of a model, plays a critical role in shaping the optimization dynamics of neural nets. While its impact has been extensively studied under the asymptotic regimes---both in training time and FLS---existing theory offers limited insight into how FLS affects generalization in practical settings, such as when training is stopped upon reaching a target training risk. In this work, we investigate the impact of FLS on generalization in deep networks under such practical conditions. Through empirical studies, we first uncover the emergence of an optimal FLS ---neither too small nor too large---that yields substantial generalization gains. This finding runs counter to the prevailing intuition that stronger feature learning universally improves generalization. To explain this phenomenon, we develop a theoretical analysis of gradient flow dynamics in two-layer ReLU nets trained with logistic loss, where FLS is controlled via initialization scale. Our main theoretical result establishes the existence of an optimal FLS arising from a trade-off between two competing effects: An excessively large FLS induces an over-alignment phenomenon that degrades generalization, while an overly small FLS leads to over-fitting .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 70b1c96c-4fa8-4186-bd65-129aee0ac26eBuilds on29
- Gradient Descent Maximizes the Margin of Homogeneous Neural NetworksKaifeng Lyu, Jian LiICLR 2020 · 402 citations
- The Intrinsic Dimension of Images and Its Impact on LearningPhillip Pope, Chen Zhu, Ahmed Abdelkader, Micah Goldblum et al.ICLR 2021 · 381 citations
- Tensor Programs IV: Feature Learning in Infinite-Width Neural NetworksGreg Yang, Edward J. HuICML 2021 · 242 citations
- Tuning Large Neural Networks via Zero-Shot Hyperparameter TransferGe Yang, Edward J. Hu, Igor Babuschkin, Szymon Sidor et al.NeurIPS 2021 · 208 citations
- Self-Consistent Dynamical Field Theory of Kernel Evolution in Wide Neural NetworksBlake Bordelon, Cengiz PehlevanNeurIPS 2022 · 140 citations
Related papers
- Dichotomy of Feature Learning and Unlearning: Fast-Slow Analysis on Neural Networks with Stochastic Gradient DescentShota Imai, Sota Nishiyama, Masaaki ImaizumiICML 2026
- Extreme Memorization via Scale of InitializationHarsh Mehta, Ashok Cutkosky, Behnam NeyshaburICLR 2021 · 22 citations
- Simplicity Bias and Optimization Threshold in Two-Layer ReLU NetworksEtienne Boursier, Nicolas FlammarionICML 2025
- How Much Over-parameterization Is Sufficient to Learn Deep ReLU Networks?Zixiang Chen, Yuan Cao, Difan Zou, Quanquan GuICLR 2021 · 29 citations
- Generalization of Two-layer Neural Networks: An Asymptotic ViewpointJimmy Ba, Murat A. Erdogdu, Taiji Suzuki, Denny Wu et al.ICLR 2020 · 77 citations
