Why Spectral Normalization Stabilizes GANs: Analysis and Improvements
Zinan Lin, Vyas Sekar, Giulia Fanti
摘要
Spectral normalization (SN) is a widely-used technique for improving the stability of Generative Adversarial Networks (GANs) by forcing each layer of the discriminator to have unit spectral norm. This approach controls the Lipschitz constant of the discriminator, and is empirically known to improve sample quality in many GAN architectures. However, there is currently little understanding of why SN is so effective. In this work, we show that SN controls two important failure modes of GAN training: exploding and vanishing gradients. Our proofs illustrate a (perhaps unintentional) connection with the successful LeCun initialization technique, proposed over two decades ago to control gradients in the training of deep neural networks. This connection helps to explain why the most popular implementation of SN for GANs requires no hyperparameter tuning, whereas stricter implementations of SN have poor empirical performance out-of-the-box. Unlike LeCun initialization which only controls gradient vanishing at the beginning of training, we show that SN tends to preserve this property throughout training. Finally, building on this theoretical understanding, we propose Bidirectional Spectral Normalization (BSN), a modification of SN inspired by Xavier initialization, a later improvement to LeCun initialization. Theoretically, we show that BSN gives better gradient control than SN. Empirically, we demonstrate that BSN outperforms SN in sample quality on several benchmark datasets, while also exhibiting better training stability.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Understanding the Generalization Benefit of Normalization Layers: Sharpness ReductionKaifeng Lyu, Zhiyuan Li, Sanjeev AroraNeurIPS 2022 · 被引用 111 次
- SNN-RAT: Robustness-enhanced Spiking Neural Network through Regularized Adversarial TrainingJianhao Ding, Tong Bu, Zhaofei Yu, Tiejun Huang 等NeurIPS 2022 · 被引用 70 次
- SAN: Inducing Metrizability of GAN with Discriminative Normalized Linear LayerYuhta Takida, Masaaki Imaizumi, Takashi Shibuya, Chieh-Hsin Lai 等ICLR 2024 · 被引用 28 次
- Fast Mixing of Stochastic Gradient Descent with Normalization and Weight DecayZhiyuan Li, Tianhao Wang, Dingli YuNeurIPS 2022 · 被引用 19 次
- Transformers Get Stable: An End-to-End Signal Propagation Theory for Language ModelsAkhil Kedia, Mohd Abbas Zaidi, Sushil Khyalia, Jungho Jung 等ICML 2024 · 被引用 16 次
它引用的顶会 Paper1
相关 Paper
- Gradient Normalization for Generative Adversarial NetworksYi-Lun Wu, Hong-Han Shuai, Zhi Rui Tam, Hong-Yu ChiuICCV 2021 · 被引用 78 次
- Controllable Orthogonalization in Training DNNsLei Huang, Li Liu, Fan Zhu, Diwen Wan 等CVPR 2020
- Spectral Regularization for Combating Mode Collapse in GANsKanglin Liu, Guoping Qiu, Wenming Tang, Fei ZhouICCV 2019 · 被引用 97 次
- Sparsity Aware Normalization for GANsIdan Kligvasser, Tomer MichaeliAAAI 2021 · 被引用 6 次
- The Impact of Neural Network Overparameterization on Gradient Confusion and Stochastic Gradient DescentKarthik Abinav Sankararaman, Soham De, Zheng Xu, W. Ronny Huang 等ICML 2020 · 被引用 122 次
