Unveiling The Matthew Effect Across Channels: Assessing Layer Width Sufficiency via Weight Norm Variance
Yiting Chen, Jiazi Bu, Junchi Yan
Abstract
The trade-off between cost and performance has been a longstanding and critical issue for deep neural networks. One key factor affecting the computational cost is the width of each layer. However, in practice, the width of layers in a neural network is mostly empirically determined. In this paper, we show that a pattern regarding the variance of weight norm corresponding to different channels can indicate whether the layer is sufficiently wide and may help us better allocate computational resources across the layers. Starting from a simple intuition that channels with larger weights would have larger gradients and the difference in weight norm enlarges between channels with similar weight, we empirically validate that wide and narrow layers show two different patterns with experiments across different data modalities and network architectures. Based on the two different patterns, we identify three stages during training and explain each stage with corresponding evidence. We further propose to adjust the width based on the identified pattern and show that conventional layer width settings for CNNs could be adjusted to reduce the number of parameters while boosting the performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on11
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang et al.ICLR 2020 · 1,108 citations
- The Early Phase of Neural Network TrainingJonathan Frankle, David J. Schwab, Ari S. MorcosICLR 2020 · 199 citations
- Zero-CL: Instance and Feature decorrelation for negative-free symmetric contrastive learningShaofeng Zhang, Feng Zhu, Junchi Yan, Rui Zhao et al.ICLR 2022 · 52 citations
- Why neural networks find simple solutions: The many regularizers of geometric complexityBenoit Dherin, Michael Munn, Mihaela Rosca, David BarrettNeurIPS 2022 · 52 citations
Related papers
- Batch normalization is sufficient for universal function approximation in CNNsRebekka BurkholzICLR 2024 · 8 citations
- The Heterogeneity Hypothesis: Finding Layer-Wise Differentiated Network ArchitecturesYawei Li, Wen Li, Martin Danelljan, Kai Zhang et al.CVPR 2021
- Which Frequencies do CNNs Need? Emergent Bottleneck Structure in Feature LearningYuxiao Wen, Arthur JacotICML 2024 · 9 citations
- Redundant representations help generalization in wide neural networksDiego Doimo, Aldo Glielmo, Sebastian Goldt, Alessandro LaioNeurIPS 2022 · 13 citations
- Are wider nets better given the same number of parameters?Anna Golubeva, Guy Gur-Ari, Behnam NeyshaburICLR 2021 · 48 citations
