Batch normalization is sufficient for universal function approximation in CNNs
Rebekka Burkholz
Abstract
Normalization techniques, for which Batch Normalization (BN) is a popular choice, is an integral part of many deep learning architectures and contributes significantly to the learning success. We provide a partial explanation for this phenomenon by proving that training normalization parameters alone is already sufficient for universal function approximation if the number of available, potentially random features matches or exceeds the weight parameters of the target networks that can be expressed. Our bound on the number of required features does not only improve on a recent result for fully-connected feed-forward architectures but also applies to CNNs with and without residual connections and almost arbitrary activation functions (which include ReLUs). Our explicit construction of a given target network solves a depth-width trade-off that is driven by architectural constraints and can explain why switching off entire neurons can have representational benefits, as has been observed empirically. To validate our theory, we explicitly match target networks that outperform experimentally obtained networks with trained BN parameters by utilizing a sufficient number of random features.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Sign-In to the Lottery: Reparameterizing Sparse TrainingAdvait Gadhikar, Tom Jacobs, Chao Zhou, Rebekka BurkholzNeurIPS 2025
- Expressivity of Neural Networks with Random Weights and Learned BiasesEzekiel Williams, Alexandre Payeur, Avery Hee-Woon Ryoo, Thomas Jiralerspong et al.ICLR 2025
- Test-time Adaptation for Regression by Subspace AlignmentKazuki Adachi, Shin'ya Yamaguchi, Atsutoshi Kumagai, Tomoki HamagamiICLR 2025
Builds on17
- High-Performance Large-Scale Image Recognition Without NormalizationAndy Brock, Soham De, Samuel L. Smith, Karen SimonyanICML 2021 · 613 citations
- Batch Normalization Biases Residual Blocks Towards the Identity Function in Deep NetworksSoham De, Samuel L. SmithNeurIPS 2020 · 173 citations
- Training BatchNorm and Only BatchNorm: On the Expressive Power of Random Features in CNNsJonathan Frankle, David J. Schwab, Ari S. MorcosICLR 2021 · 163 citations
- The Unreasonable Effectiveness of Random Pruning: Return of the Most Naive Baseline for Sparse TrainingShiwei Liu, Tianlong Chen, Xiaohan Chen, Li Shen et al.ICLR 2022 · 141 citations
- Frozen Pretrained Transformers as Universal Computation EnginesKevin Lu, Aditya Grover, Pieter Abbeel, Igor MordatchAAAI 2022 · 133 citations
Related papers
- Most Activation Functions Can Win the Lottery Without Excessive DepthRebekka BurkholzNeurIPS 2022 · 27 citations
- Deep Network Approximation in Terms of Intrinsic ParametersZuowei Shen, Haizhao Yang, Shijun ZhangICML 2022 · 13 citations
- Beyond BatchNorm: Towards a Unified Understanding of Normalization in Deep LearningEkdeep Singh Lubana, Robert P. Dick, Hidenori TanakaNeurIPS 2021 · 50 citations
- A Probabilistic Approach to Neural Network PruningXin Qian, Diego KlabjanICML 2021 · 24 citations
- Compelling ReLU Networks to Exhibit Exponentially Many Linear Regions at Initialization and During TrainingMax Milkert, David Hyde, Forrest J. LaineICML 2025
