Ghost Noise for Regularizing Deep Neural Networks
Atli Kosson, Dongyang Fan, Martin Jaggi
Abstract
Batch Normalization (BN) is widely used to stabilize the optimization process and improve the test performance of deep neural networks. The regularization effect of BN depends on the batch size and explicitly using smaller batch sizes with Batch Normalization, a method known as Ghost Batch Normalization (GBN), has been found to improve generalization in many settings. We investigate the effectiveness of GBN by disentangling the induced ``Ghost Noise'' from normalization and quantitatively analyzing the distribution of noise as well as its impact on model performance. Inspired by our analysis, we propose a new regularization technique called Ghost Noise Injection (GNI) that imitates the noise in GBN without incurring the detrimental train-test discrepancy effects of small batch training. We experimentally show that GNI can provide a greater generalization benefit than GBN. Ghost Noise Injection can also be beneficial in otherwise non-noisy settings such as layer-normalized networks, providing additional evidence of the usefulness of Ghost Noise in Batch Normalization as a regularizer.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 36cf7a2c-e5e0-42d9-8fba-1c16360edf85Cited by top-tier papers2
- Rotational Equilibrium: How Weight Decay Balances Learning Across Neural NetworksAtli Kosson, Bettina Messmer, Martin JaggiICML 2024 · 39 citations
- Stable Coresets via Posterior Sampling: Aligning Induced and Full Loss LandscapesWei-Kai Chang, Rajiv KhannaNeurIPS 2025
Builds on7
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- On the Relationship between Self-Attention and Convolutional LayersJean-Baptiste Cordonnier, Andreas Loukas, Martin JaggiICLR 2020 · 629 citations
- High-Performance Large-Scale Image Recognition Without NormalizationAndy Brock, Soham De, Samuel L. Smith, Karen SimonyanICML 2021 · 613 citations
- The Implicit and Explicit Regularization Effects of DropoutColin Wei, Sham M. Kakade, Tengyu MaICML 2020 · 129 citations
- Explicit Regularisation in Gaussian Noise InjectionsAlexander Camuto, Matthew Willetts, Umut Simsekli, Stephen J. Roberts et al.NeurIPS 2020 · 90 citations
Related papers
- Four Things Everyone Should Know to Improve Batch NormalizationCecilia Summers, Michael J. DinneenICLR 2020 · 57 citations
- Instance Enhancement Batch Normalization: An Adaptive Regulator of Batch NoiseSenwei Liang, Zhongzhan Huang, Mingfu Liang, Haizhao YangAAAI 2020 · 65 citations
- An Investigation Into the Stochasticity of Batch WhiteningLei Huang, Lei Zhao, Yi Zhou, Fan Zhu et al.CVPR 2020
- Noise Injection Node Regularization for Robust LearningNoam Levi, Itay M. Bloch, Marat Freytsis, Tomer VolanskyICLR 2023 · 2 citations
- Understanding the Failure of Batch Normalization for Transformers in NLPJiaxi Wang, Ji Wu, Lei HuangNeurIPS 2022 · 14 citations
