Initial Guessing Bias: How Untrained Networks Favor Some Classes
Emanuele Francazi, Aurélien Lucchi, Marco Baity-Jesi
Abstract
Understanding and controlling biasing effects in neural networks is crucial for ensuring accurate and fair model performance. In the context of classification problems, we provide a theoretical analysis demonstrating that the structure of a deep neural network (DNN) can condition the model to assign all predictions to the same class, even before the beginning of training, and in the absence of explicit biases. We prove that, besides dataset properties, the presence of this phenomenon, which we call Initial Guessing Bias (IGB), is influenced by model choices including dataset preprocessing methods, and architectural decisions, such as activation functions, max-pooling layers, and network depth. Our analysis of IGB provides information for architecture selection and model initialization. We also highlight theoretical consequences, such as the breakdown of node-permutation symmetry, the violation of self-averaging and the non-trivial effects that depth has on the phenomenon.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c2b759d6-0ab3-400f-93c2-b2de5b682b42Cited by top-tier papers3
- Bias in Motion: Theoretical Insights into the Dynamics of Bias in SGD TrainingAnchit Jain, Rozhin Nobahari, Aristide Baratin, Stefano Sarao MannelliNeurIPS 2024 · 8 citations
- Neural Redshift: Random Networks are not Random FunctionsDamien Teney, Armand Mihai Nicolicioiu, Valentin Hartmann, Ehsan AbbasnejadCVPR 2024 · 7 citations
- When Bias Meets Trainability: Connecting Theories of InitializationAlberto Bassi, Marco Baity-Jesi, Aurélien Lucchi, Carlo Albert et al.ICLR 2026
Builds on6
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer et al.NeurIPS 2021 · 3,862 citations
- Understanding the Dynamics of Gradient Flow in Overparameterized Linear modelsSalma Tarmoun, Guilherme França, Benjamin D. Haeffele, René VidalICML 2021 · 76 citations
- A Theoretical Analysis of the Learning Dynamics under Class ImbalanceEmanuele Francazi, Marco Baity-Jesi, Aurélien LucchiICML 2023 · 33 citations
- How much does Initialization Affect Generalization?Sameera Ramasinghe, Lachlan Ewen MacDonald, Moshiur R. Farazi, Hemanth Saratchandran et al.ICML 2023 · 9 citations
Related papers
- Beyond Signal Propagation: Is Feature Diversity Necessary in Deep Neural Network Initialization?Yaniv Blumenfeld, Dar Gilboa, Daniel SoudryICML 2020 · 18 citations
- Simplicity Bias in Overparameterized Machine LearningYakir BerchenkoAAAI 2024 · 7 citations
- The future is log-Gaussian: ResNets and their infinite-depth-and-width limit at initializationMufan Bill Li, Mihai Nica, Daniel M. RoyNeurIPS 2021 · 41 citations
- Let's Agree to Agree: Neural Networks Share Classification Order on Real DatasetsGuy Hacohen, Leshem Choshen, Daphna WeinshallICML 2020 · 63 citations
- Neural networks trained with SGD learn distributions of increasing complexityMaria Refinetti, Alessandro Ingrosso, Sebastian GoldtICML 2023 · 58 citations
