Beyond Signal Propagation: Is Feature Diversity Necessary in Deep Neural Network Initialization?
Yaniv Blumenfeld, Dar Gilboa, Daniel Soudry
Abstract
Deep neural networks are typically initialized with random weights, with variances chosen to facilitate signal propagation and stable gradients. It is also believed that diversity of features is an important property of these initializations. We construct a deep convolutional network with identical features by initializing almost all the weights to . The architecture also enables perfect signal propagation and stable gradients, and achieves high accuracy on standard benchmarks. This indicates that random, diverse initializations are not necessary for training neural networks. An essential element in training this network is a mechanism of symmetry breaking; we study this phenomenon and find that standard GPU operations, which are non-deterministic, can serve as a sufficient source of symmetry breaking to enable training.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 26451b3d-1ad0-40bd-9d03-408ceed43ff5Cited by top-tier papers6
- The future is log-Gaussian: ResNets and their infinite-depth-and-width limit at initializationMufan Bill Li, Mihai Nica, Daniel M. RoyNeurIPS 2021 · 41 citations
- Parseval Regularization for Continual Reinforcement LearningWesley Chung, Lynn Cherif, Doina Precup, David MegerNeurIPS 2024 · 27 citations
- Learning on a Razor's Edge: Identifiability and Singularity of Polynomial Neural NetworksVahid Shahverdi, Giovanni Luca Marchetti, Kathlén KohnICLR 2026 · 11 citations
- Sinusoidal Initialization, Time for a New StartAlberto Fernández-Hernández, José I. Mestre, Manuel F. Dolz, José Duato et al.NeurIPS 2025 · 5 citations
- Leveraging Sub-Optimal Data for Human-in-the-Loop Reinforcement LearningCalarina Muslimani, Matthew E. TaylorICLR 2025
Builds on1
Related papers
- Deep Isometric Learning for Visual RecognitionHaozhi Qi, Chong You, Xiaolong Wang, Yi Ma et al.ICML 2020 · 57 citations
- Deep equilibrium networks are sensitive to initialization statisticsAtish Agarwala, Samuel S. SchoenholzICML 2022 · 12 citations
- Initial Guessing Bias: How Untrained Networks Favor Some ClassesEmanuele Francazi, Aurélien Lucchi, Marco Baity-JesiICML 2024 · 7 citations
- IDInit: A Universal and Stable Initialization Method for Neural Network TrainingYu Pan, Chaozheng Wang, Zekai Wu, Qifan Wang et al.ICLR 2025
- Understanding the Covariance Structure of Convolutional FiltersAsher Trockman, Devin Willmott, J. Zico KolterICLR 2023
