Training BatchNorm and Only BatchNorm: On the Expressive Power of Random Features in CNNs
Jonathan Frankle, David J. Schwab, Ari S. Morcos
Abstract
A wide variety of deep learning techniques from style transfer to multitask learning rely on training affine transformations of features. Most prominent among these is the popular feature normalization technique BatchNorm, which normalizes activations and then subsequently applies a learned affine transform. In this paper, we aim to understand the role and expressive power of affine parameters used to transform features in this way. To isolate the contribution of these parameters from that of the learned features they transform, we investigate the performance achieved when training only these parameters in BatchNorm and freezing all weights at their random initializations. Doing so leads to surprisingly high performance considering the significant limitations that this style of training imposes. For example, sufficiently deep ResNets reach 82% (CIFAR-10) and 32% (ImageNet, top-5) accuracy in this configuration, far higher than when training an equivalent number of randomly chosen parameters elsewhere in the network. BatchNorm achieves this performance in part by naturally learning to disable around a third of the random features. Not only do these results highlight the expressive power of affine parameters in deep learning, but-in a broader sense-they characterize the expressive power of neural networks constructed simply by shifting and rescaling random features.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 59e786d3-b6c7-4b1a-90a4-a467c8680030Cited by top-tier papers37
- Improving robustness against common corruptions by covariate shift adaptationSteffen Schneider, Evgenia Rusak, Luisa Eck, Oliver Bringmann et al.NeurIPS 2020 · 688 citations
- TinyTL: Reduce Memory, Not Parameters for Efficient On-Device LearningHan Cai, Chuang Gan, Ligeng Zhu, Song HanNeurIPS 2020 · 375 citations
- Supermasks in SuperpositionMitchell Wortsman, Vivek Ramanujan, Rosanne Liu, Aniruddha Kembhavi et al.NeurIPS 2020 · 364 citations
- On-Device Training Under 256KB MemoryJi Lin, Ligeng Zhu, Wei-Ming Chen, Wei-Chen Wang et al.NeurIPS 2022 · 345 citations
- VeRA: Vector-based Random Matrix AdaptationDawid Jan Kopiczko, Tijmen Blankevoort, Yuki M. AsanoICLR 2024 · 308 citations
Builds on1
Related papers
- Batch normalization is sufficient for universal function approximation in CNNsRebekka BurkholzICLR 2024 · 8 citations
- Revisiting Learnable Affines for Batch Norm in Few-Shot Transfer LearningMoslem Yazdanpanah, Aamer Abdul Rahman, Muawiz Chaudhary, Christian Desrosiers et al.CVPR 2022 · 18 citations
- Beyond BatchNorm: Towards a Unified Understanding of Normalization in Deep LearningEkdeep Singh Lubana, Robert P. Dick, Hidenori TanakaNeurIPS 2021 · 50 citations
- Deconstructing the Regularization of BatchNormYann N. Dauphin, Ekin Dogus CubukICLR 2021 · 6 citations
- Normalization Layers Are All That Sharpness-Aware Minimization NeedsMaximilian Müller, Tiffany Vlaar, David Rolnick, Matthias HeinNeurIPS 2023 · 37 citations
