Investigating Generalization by Controlling Normalized Margin
Alexander R. Farhang, Jeremy D. Bernstein, Kushal Tirumala, Yang Liu, Yisong Yue
Abstract
Weight norm (cid:107) w (cid:107) and margin γ participate in learning theory via the normalized margin γ/ (cid:107) w (cid:107) . Since standard neural net optimizers do not control normalized margin, it is hard to test whether this quantity causally relates to generalization. This paper designs a series of experimental studies that explicitly control normalized margin and thereby tackle two central questions. First: does normalized margin always have a causal effect on generalization? The paper finds that no — networks can be produced where normalized margin has seemingly no relationship with generalization, counter to the theory of Bartlett et al. (2017). Second: does normalized margin ever have a causal effect on generalization? The paper finds that yes —in a standard training setup, test performance closely tracks normalized margin. The paper suggests a Gaussian process model as a promising explanation for this behavior.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 551cf519-ce89-4f24-a4d1-ccc706beac40Cited by top-tier papers1
Ask how each one uses itBuilds on7
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang et al.ICLR 2020 · 1,108 citations
- Fantastic Generalization Measures and Where to Find ThemYiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan et al.ICLR 2020 · 705 citations
- Evaluation of Neural Architectures trained with square Loss vs Cross-Entropy in Classification TasksLike Hui, Mikhail BelkinICLR 2021 · 199 citations
- In search of robust measures of generalizationGintare Karolina Dziugaite, Alexandre Drouin, Brady Neal, Nitarshan Rajkumar et al.NeurIPS 2020 · 112 citations
- PAC-Bayes Analysis Beyond the Usual BoundsOmar Rivasplata, Ilja Kuzborskij, Csaba Szepesvári, John Shawe-TaylorNeurIPS 2020 · 101 citations
Related papers
- Improved Sample Complexities for Deep Neural Networks and Robust Classification via an All-Layer MarginColin Wei, Tengyu MaICLR 2020 · 91 citations
- Understanding the Generalization Benefit of Normalization Layers: Sharpness ReductionKaifeng Lyu, Zhiyuan Li, Sanjeev AroraNeurIPS 2022 · 111 citations
- Formalizing Generalization and Adversarial Robustness of Neural Networks to Weight PerturbationsYu-Lin Tsai, Chia-Yi Hsu, Chia-Mu Yu, Pin-Yu ChenNeurIPS 2021 · 36 citations
- Optimization Theory for ReLU Neural Networks Trained with Normalization LayersYonatan Dukler, Quanquan Gu, Guido MontúfarICML 2020 · 30 citations
- Input Margins Can Predict Generalization TooCoenraad Mouton, Marthinus Wilhelmus Theunissen, Marelie H. DavelAAAI 2024 · 5 citations
