Dropout Universality: Scaling Laws and Optimal Scheduling at the Edge-of-Chaos
Lucas Fernandez-Sarmiento
Abstract
We develop a mean-field theory of dropout as a perturbation of critical signal propagation at the edge of chaos, and show that it predicts a simple, no-cost change to standard practice: front-loaded dropout schedules cut test loss by 18–35% over constant dropout in MLPs and Vision Transformers at fixed budget. The theoretical mechanism is that dropout shifts the perfect-alignment fixed point, making the depth scale for information propagation finite even at critical initialization. We derive critical and crossover scaling laws for correlation decay and establish that smooth activations and kinked, ReLU-like activations constitute distinct universality classes, with different critical exponents and a universal two-parameter scaling collapse in detuning and dropout strength. The distinction traces to the analytic structure of the correlation map: smooth activations admit a Taylor expansion near perfect alignment, while kinked activations develop a branch point with universal non-analyticity. As a corollary, the framework yields saturated dropout profiles under fixed budget; a regularization-reach argument then selects front-loaded schedules, with accuracy gains as a consistent secondary effect. We also discuss how the same Gaussian-kernel structure extends the theory beyond MLPs toward CNNs and residual architectures.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on4
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Reducing Transformer Depth on Demand with Structured DropoutAngela Fan, Edouard Grave, Armand JoulinICLR 2020 · 695 citations
- Infinite attention: NNGP and NTK for deep attention networksJiri Hron, Yasaman Bahri, Jascha Sohl-Dickstein, Roman NovakICML 2020 · 147 citations
- Quadratic models for understanding catapult dynamics of neural networksLibin Zhu, Chaoyue Liu, Adityanarayanan Radhakrishnan, Mikhail BelkinICLR 2024 · 17 citations
Related papers
- Don't Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and InferenceMostafa Elhoushi, Alexander Pretko, Nolan Dey, Bin Zhang et al.ICML 2026
- The phase diagram of approximation rates for deep neural networksDmitry Yarotsky, Anton ZhevnerchukNeurIPS 2020 · 156 citations
- Critical Initialization of Wide and Deep Neural Networks using Partial Jacobians: General Theory and ApplicationsDarshil Doshi, Tianyu He, Andrey GromovNeurIPS 2023 · 10 citations
- The future is log-Gaussian: ResNets and their infinite-depth-and-width limit at initializationMufan Bill Li, Mihai Nica, Daniel M. RoyNeurIPS 2021 · 41 citations
- Beyond ReLU: Bifurcation, Oversmoothing, and Topological PriorsErkan Turan, Gaspard Abel, Maysam Behmanesh, Emery Pierson et al.ICML 2026 · 1 citation
