Why Spectral Normalization Stabilizes GANs: Analysis and Improvements
Zinan Lin, Vyas Sekar, Giulia Fanti
Abstract
Spectral normalization (SN) is a widely-used technique for improving the stability of Generative Adversarial Networks (GANs) by forcing each layer of the discriminator to have unit spectral norm. This approach controls the Lipschitz constant of the discriminator, and is empirically known to improve sample quality in many GAN architectures. However, there is currently little understanding of why SN is so effective. In this work, we show that SN controls two important failure modes of GAN training: exploding and vanishing gradients. Our proofs illustrate a (perhaps unintentional) connection with the successful LeCun initialization technique, proposed over two decades ago to control gradients in the training of deep neural networks. This connection helps to explain why the most popular implementation of SN for GANs requires no hyperparameter tuning, whereas stricter implementations of SN have poor empirical performance out-of-the-box. Unlike LeCun initialization which only controls gradient vanishing at the beginning of training, we show that SN tends to preserve this property throughout training. Finally, building on this theoretical understanding, we propose Bidirectional Spectral Normalization (BSN), a modification of SN inspired by Xavier initialization, a later improvement to LeCun initialization. Theoretically, we show that BSN gives better gradient control than SN. Empirically, we demonstrate that BSN outperforms SN in sample quality on several benchmark datasets, while also exhibiting better training stability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fb3498b1-ae73-4bcc-8042-2dd2c9dc53ceCited by top-tier papers11
- Understanding the Generalization Benefit of Normalization Layers: Sharpness ReductionKaifeng Lyu, Zhiyuan Li, Sanjeev AroraNeurIPS 2022 · 111 citations
- SNN-RAT: Robustness-enhanced Spiking Neural Network through Regularized Adversarial TrainingJianhao Ding, Tong Bu, Zhaofei Yu, Tiejun Huang et al.NeurIPS 2022 · 70 citations
- SAN: Inducing Metrizability of GAN with Discriminative Normalized Linear LayerYuhta Takida, Masaaki Imaizumi, Takashi Shibuya, Chieh-Hsin Lai et al.ICLR 2024 · 28 citations
- Fast Mixing of Stochastic Gradient Descent with Normalization and Weight DecayZhiyuan Li, Tianhao Wang, Dingli YuNeurIPS 2022 · 19 citations
- Transformers Get Stable: An End-to-End Signal Propagation Theory for Language ModelsAkhil Kedia, Mohd Abbas Zaidi, Sushil Khyalia, Jungho Jung et al.ICML 2024 · 16 citations
Builds on1
Related papers
- Gradient Normalization for Generative Adversarial NetworksYi-Lun Wu, Hong-Han Shuai, Zhi Rui Tam, Hong-Yu ChiuICCV 2021 · 78 citations
- Controllable Orthogonalization in Training DNNsLei Huang, Li Liu, Fan Zhu, Diwen Wan et al.CVPR 2020
- Spectral Regularization for Combating Mode Collapse in GANsKanglin Liu, Guoping Qiu, Wenming Tang, Fei ZhouICCV 2019 · 97 citations
- Sparsity Aware Normalization for GANsIdan Kligvasser, Tomer MichaeliAAAI 2021 · 6 citations
- The Impact of Neural Network Overparameterization on Gradient Confusion and Stochastic Gradient DescentKarthik Abinav Sankararaman, Soham De, Zheng Xu, W. Ronny Huang et al.ICML 2020 · 122 citations
