Training BatchNorm and Only BatchNorm: On the Expressive Power of Random Features in CNNs
Jonathan Frankle, David J. Schwab, Ari S. Morcos
摘要
A wide variety of deep learning techniques from style transfer to multitask learning rely on training affine transformations of features. Most prominent among these is the popular feature normalization technique BatchNorm, which normalizes activations and then subsequently applies a learned affine transform. In this paper, we aim to understand the role and expressive power of affine parameters used to transform features in this way. To isolate the contribution of these parameters from that of the learned features they transform, we investigate the performance achieved when training only these parameters in BatchNorm and freezing all weights at their random initializations. Doing so leads to surprisingly high performance considering the significant limitations that this style of training imposes. For example, sufficiently deep ResNets reach 82% (CIFAR-10) and 32% (ImageNet, top-5) accuracy in this configuration, far higher than when training an equivalent number of randomly chosen parameters elsewhere in the network. BatchNorm achieves this performance in part by naturally learning to disable around a third of the random features. Not only do these results highlight the expressive power of affine parameters in deep learning, but-in a broader sense-they characterize the expressive power of neural networks constructed simply by shifting and rescaling random features.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper37
- Improving robustness against common corruptions by covariate shift adaptationSteffen Schneider, Evgenia Rusak, Luisa Eck, Oliver Bringmann 等NeurIPS 2020 · 被引用 688 次
- TinyTL: Reduce Memory, Not Parameters for Efficient On-Device LearningHan Cai, Chuang Gan, Ligeng Zhu, Song HanNeurIPS 2020 · 被引用 375 次
- Supermasks in SuperpositionMitchell Wortsman, Vivek Ramanujan, Rosanne Liu, Aniruddha Kembhavi 等NeurIPS 2020 · 被引用 364 次
- On-Device Training Under 256KB MemoryJi Lin, Ligeng Zhu, Wei-Ming Chen, Wei-Chen Wang 等NeurIPS 2022 · 被引用 345 次
- VeRA: Vector-based Random Matrix AdaptationDawid Jan Kopiczko, Tijmen Blankevoort, Yuki M. AsanoICLR 2024 · 被引用 308 次
它引用的顶会 Paper1
相关 Paper
- Batch normalization is sufficient for universal function approximation in CNNsRebekka BurkholzICLR 2024 · 被引用 8 次
- Revisiting Learnable Affines for Batch Norm in Few-Shot Transfer LearningMoslem Yazdanpanah, Aamer Abdul Rahman, Muawiz Chaudhary, Christian Desrosiers 等CVPR 2022 · 被引用 18 次
- Beyond BatchNorm: Towards a Unified Understanding of Normalization in Deep LearningEkdeep Singh Lubana, Robert P. Dick, Hidenori TanakaNeurIPS 2021 · 被引用 50 次
- Deconstructing the Regularization of BatchNormYann N. Dauphin, Ekin Dogus CubukICLR 2021 · 被引用 6 次
- Normalization Layers Are All That Sharpness-Aware Minimization NeedsMaximilian Müller, Tiffany Vlaar, David Rolnick, Matthias HeinNeurIPS 2023 · 被引用 37 次
