Stable Rank Normalization for Improved Generalization in Neural Networks and GANs
Amartya Sanyal, Philip H. S. Torr, Puneet K. Dokania
摘要
Exciting new work on the generalization bounds for neural networks (NN) given by Neyshabur et al. , Bartlett et al. closely depend on two parameter-depenedent quantities: the Lipschitz constant upper-bound and the stable rank (a softer version of the rank operator). This leads to an interesting question of whether controlling these quantities might improve the generalization behaviour of NNs. To this end, we propose stable rank normalization (SRN), a novel, optimal, and computationally efficient weight-normalization scheme which minimizes the stable rank of a linear operator. Surprisingly we find that SRN, inspite of being non-convex problem, can be shown to have a unique optimal solution. Moreover, we show that SRN allows control of the data-dependent empirical Lipschitz constant, which in contrast to the Lipschitz upper-bound, reflects the true behaviour of a model on a given dataset. We provide thorough analyses to show that SRN, when applied to the linear layers of a NN for classification, provides striking improvements-11.3% on the generalization gap compared to the standard NN along with significant reduction in memorization. When applied to the discriminator of GANs (called SRN-GAN) it improves Inception, FID, and Neural divergence scores on the CIFAR 10/100 and CelebA datasets, while learning mappings with low empirical Lipschitz constants.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- Using Mixup as a Regularizer Can Surprisingly Improve Accuracy & Out-of-Distribution RobustnessFrancesco Pinto, Harry Yang, Ser Nam Lim, Philip H. S. Torr 等NeurIPS 2022 · 被引用 74 次
- What Makes and Breaks Safety Fine-tuning? A Mechanistic StudySamyak Jain, Ekdeep Singh Lubana, Kemal Oksuz, Tom Joy 等NeurIPS 2024 · 被引用 62 次
- How Benign is Benign Overfitting ?Amartya Sanyal, Puneet K. Dokania, Varun Kanade, Philip H. S. TorrICLR 2021 · 被引用 61 次
- Improving GAN Training with Probability Ratio Clipping and Sample ReweightingYue Wu, Pan Zhou, Andrew Gordon Wilson, Eric P. Xing 等NeurIPS 2020 · 被引用 39 次
- Stable Long-Term Recurrent Video Super-ResolutionBenjamin Naoto Chiche, Arnaud Woiselle, Joana Frontera-Pons, Jean-Luc StarckCVPR 2022 · 被引用 30 次
相关 Paper
- Gradient Normalization for Generative Adversarial NetworksYi-Lun Wu, Hong-Han Shuai, Zhi Rui Tam, Hong-Yu ChiuICCV 2021 · 被引用 78 次
- Controllable Orthogonalization in Training DNNsLei Huang, Li Liu, Fan Zhu, Diwen Wan 等CVPR 2020
- Width Independent Bounds for the Local Lipschitz Constant of Deep Neural Networks at Random Initialization and after Lazy TrainingApostolos Evangelidis, Felix KrahmerICML 2026
- Batch normalization provably avoids ranks collapse for randomly initialised deep networksHadi Daneshmand, Jonas Moritz Kohler, Francis R. Bach, Thomas Hofmann 等NeurIPS 2020 · 被引用 73 次
- Why Spectral Normalization Stabilizes GANs: Analysis and ImprovementsZinan Lin, Vyas Sekar, Giulia FantiNeurIPS 2021 · 被引用 67 次
