On the Importance of Gaussianizing Representations
Daniel Eftekhari, Vardan Papyan
Abstract
The normal distribution plays a central role in information theory -it is at the same time the best-case signal and worst-case noise distribution, has the greatest representational capacity of any distribution, and offers an equivalence between uncorrelatedness and independence for joint distributions. Accounting for the mean and variance of activations throughout the layers of deep neural networks has had a significant effect on facilitating their effective training, but seldom has a prescription for precisely what distribution these activations should take, and how this might be achieved, been offered. Motivated by the information-theoretic properties of the normal distribution, we address this question and concurrently present normality normalization: a novel normalization layer which encourages normality in the feature representations of neural networks using the power transform and employs additive Gaussian noise during training. Our experiments comprehensively demonstrate the effectiveness of normality normalization, in regards to its generalization performance on an array of widely used model and dataset combinations, its strong performance across various common factors of variation such as model width, depth, and training minibatch size, its suitability for usage wherever existing normalization layers are conventionally used, and as a means to improving model robustness to random perturbations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- InfoNCE Induces Gaussian DistributionRoy Betser, Eyal Gofer, Meir Yossef Levi, Guy GilboaICLR 2026 · 17 citations
- On Optimal Steering to Achieve Exact FairnessMohit Sharma, Amit Deshpande, Chiranjib Bhattacharyya, Rajiv Ratn ShahNeurIPS 2025 · 2 citations
Builds on4
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Switchable Whitening for Deep Representation LearningXingang Pan, Xiaohang Zhan, Jianping Shi, Xiaoou Tang et al.ICCV 2019 · 204 citations
- Batch Normalization Orthogonalizes Representations in Deep Random NetworksHadi Daneshmand, Amir Joudaki, Francis R. BachNeurIPS 2021 · 47 citations
- On Bridging the Gap between Mean Field and Finite Width Deep Random Multilayer Perceptron with Batch NormalizationAmir Joudaki, Hadi Daneshmand, Francis R. BachICML 2023 · 4 citations
Related papers
- Batch normalization is sufficient for universal function approximation in CNNsRebekka BurkholzICLR 2024 · 8 citations
- Convolutional Normalization: Improving Deep Convolutional Network Robustness and TrainingSheng Liu, Xiao Li, Yuexiang Zhai, Chong You et al.NeurIPS 2021 · 30 citations
- Encoding Robustness to Image Style via Adversarial Feature PerturbationsManli Shu, Zuxuan Wu, Micah Goldblum, Tom GoldsteinNeurIPS 2021 · 23 citations
- LayerAct: Advanced Activation Mechanism for Robust Inference of CNNsKihyuk Yoon, Chiehyeon LimAAAI 2025 · 1 citation
- Training BatchNorm and Only BatchNorm: On the Expressive Power of Random Features in CNNsJonathan Frankle, David J. Schwab, Ari S. MorcosICLR 2021 · 163 citations
