Investigating Generalization by Controlling Normalized Margin
Alexander R. Farhang, Jeremy D. Bernstein, Kushal Tirumala, Yang Liu, Yisong Yue
摘要
Weight norm (cid:107) w (cid:107) and margin γ participate in learning theory via the normalized margin γ/ (cid:107) w (cid:107) . Since standard neural net optimizers do not control normalized margin, it is hard to test whether this quantity causally relates to generalization. This paper designs a series of experimental studies that explicitly control normalized margin and thereby tackle two central questions. First: does normalized margin always have a causal effect on generalization? The paper finds that no — networks can be produced where normalized margin has seemingly no relationship with generalization, counter to the theory of Bartlett et al. (2017). Second: does normalized margin ever have a causal effect on generalization? The paper finds that yes —in a standard training setup, test performance closely tracks normalized margin. The paper suggests a Gaussian process model as a promising explanation for this behavior.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper7
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang 等ICLR 2020 · 被引用 1,108 次
- Fantastic Generalization Measures and Where to Find ThemYiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan 等ICLR 2020 · 被引用 705 次
- Evaluation of Neural Architectures trained with square Loss vs Cross-Entropy in Classification TasksLike Hui, Mikhail BelkinICLR 2021 · 被引用 199 次
- In search of robust measures of generalizationGintare Karolina Dziugaite, Alexandre Drouin, Brady Neal, Nitarshan Rajkumar 等NeurIPS 2020 · 被引用 112 次
- PAC-Bayes Analysis Beyond the Usual BoundsOmar Rivasplata, Ilja Kuzborskij, Csaba Szepesvári, John Shawe-TaylorNeurIPS 2020 · 被引用 101 次
相关 Paper
- Improved Sample Complexities for Deep Neural Networks and Robust Classification via an All-Layer MarginColin Wei, Tengyu MaICLR 2020 · 被引用 91 次
- Understanding the Generalization Benefit of Normalization Layers: Sharpness ReductionKaifeng Lyu, Zhiyuan Li, Sanjeev AroraNeurIPS 2022 · 被引用 111 次
- Formalizing Generalization and Adversarial Robustness of Neural Networks to Weight PerturbationsYu-Lin Tsai, Chia-Yi Hsu, Chia-Mu Yu, Pin-Yu ChenNeurIPS 2021 · 被引用 36 次
- Optimization Theory for ReLU Neural Networks Trained with Normalization LayersYonatan Dukler, Quanquan Gu, Guido MontúfarICML 2020 · 被引用 30 次
- Input Margins Can Predict Generalization TooCoenraad Mouton, Marthinus Wilhelmus Theunissen, Marelie H. DavelAAAI 2024 · 被引用 5 次
