Beyond Signal Propagation: Is Feature Diversity Necessary in Deep Neural Network Initialization?
Yaniv Blumenfeld, Dar Gilboa, Daniel Soudry
摘要
Deep neural networks are typically initialized with random weights, with variances chosen to facilitate signal propagation and stable gradients. It is also believed that diversity of features is an important property of these initializations. We construct a deep convolutional network with identical features by initializing almost all the weights to . The architecture also enables perfect signal propagation and stable gradients, and achieves high accuracy on standard benchmarks. This indicates that random, diverse initializations are not necessary for training neural networks. An essential element in training this network is a mechanism of symmetry breaking; we study this phenomenon and find that standard GPU operations, which are non-deterministic, can serve as a sufficient source of symmetry breaking to enable training.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- The future is log-Gaussian: ResNets and their infinite-depth-and-width limit at initializationMufan Bill Li, Mihai Nica, Daniel M. RoyNeurIPS 2021 · 被引用 41 次
- Parseval Regularization for Continual Reinforcement LearningWesley Chung, Lynn Cherif, Doina Precup, David MegerNeurIPS 2024 · 被引用 27 次
- Learning on a Razor's Edge: Identifiability and Singularity of Polynomial Neural NetworksVahid Shahverdi, Giovanni Luca Marchetti, Kathlén KohnICLR 2026 · 被引用 11 次
- Sinusoidal Initialization, Time for a New StartAlberto Fernández-Hernández, José I. Mestre, Manuel F. Dolz, José Duato 等NeurIPS 2025 · 被引用 5 次
- Leveraging Sub-Optimal Data for Human-in-the-Loop Reinforcement LearningCalarina Muslimani, Matthew E. TaylorICLR 2025
它引用的顶会 Paper1
相关 Paper
- Deep Isometric Learning for Visual RecognitionHaozhi Qi, Chong You, Xiaolong Wang, Yi Ma 等ICML 2020 · 被引用 57 次
- Deep equilibrium networks are sensitive to initialization statisticsAtish Agarwala, Samuel S. SchoenholzICML 2022 · 被引用 12 次
- Initial Guessing Bias: How Untrained Networks Favor Some ClassesEmanuele Francazi, Aurélien Lucchi, Marco Baity-JesiICML 2024 · 被引用 7 次
- IDInit: A Universal and Stable Initialization Method for Neural Network TrainingYu Pan, Chaozheng Wang, Zekai Wu, Qifan Wang 等ICLR 2025
- Understanding the Covariance Structure of Convolutional FiltersAsher Trockman, Devin Willmott, J. Zico KolterICLR 2023
