On the Nonlinearity of Layer Normalization
Yunhao Ni, Yuxin Guo, Junlong Jia, Lei Huang
Abstract
Layer normalization (LN) is a ubiquitous technique in deep learning but our theoretical understanding to it remains elusive. This paper investigates a new theoretical direction for LN, regarding to its nonlinearity and representation capacity. We investigate the representation capacity of a network with layerwise composition of linear and LN transformations, referred to as LN-Net. We theoretically show that, given samples with any label assignment, an LN-Net with only 3 neurons in each layer and LN layers can correctly classify them. We further show the lower bound of the VC dimension of an LN-Net. The nonlinearity of LN can be amplified by group partition, which is also theoretically demonstrated with mild assumption and empirically supported by our experiments. Based on our analyses, we consider to design neural architecture by exploiting and amplifying the nonlinearity of LN, and the effectiveness is supported by our experiments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8dc15c2c-49f8-434e-a2ef-c8660cac3ec4Cited by top-tier papers1
Ask how each one uses itBuilds on12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- On Layer Normalization in the Transformer ArchitectureRuibin Xiong, Yunchang Yang, Di He, Kai Zheng et al.ICML 2020 · 1,388 citations
Related papers
- On the VC dimension of deep group convolutional neural networksAnna Sepliarskaia, Sophie Langer, Johannes Schmidt-HieberNeurIPS 2025 · 1 citation
- Batch normalization is sufficient for universal function approximation in CNNsRebekka BurkholzICLR 2024 · 8 citations
- Group Whitening: Balancing Learning Efficiency and Representational CapacityLei Huang, Yi Zhou, Li Liu, Fan Zhu et al.CVPR 2021
- Beyond BatchNorm: Towards a Unified Understanding of Normalization in Deep LearningEkdeep Singh Lubana, Robert P. Dick, Hidenori TanakaNeurIPS 2021 · 50 citations
- On the impact of activation and normalization in obtaining isometric embeddings at initializationAmir Joudaki, Hadi Daneshmand, Francis R. BachNeurIPS 2023 · 16 citations
