Scaling Properties of Deep Residual Networks
Alain-Sam Cohen, Rama Cont, Alain Rossier, Renyuan Xu
Abstract
Residual networks (ResNets) have displayed impressive results in pattern recognition and, recently, have garnered considerable theoretical interest due to a perceived link with neural ordinary differential equations (neural ODEs). This link relies on the convergence of network weights to a smooth function as the number of layers increases. We investigate the properties of weights trained by stochastic gradient descent and their scaling with network depth through detailed numerical experiments. We observe the existence of scaling regimes markedly different from those assumed in neural ODE literature. Depending on certain features of the network architecture, such as the smoothness of the activation function, one may obtain an alternative ODE limit, a stochastic differential equation or neither of these. These findings cast doubts on the validity of the neural ODE model as an adequate asymptotic description of deep ResNets and point to an alternative class of differential equations as a better description of the deep network limit.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6e60c4a4-6fef-4473-8693-09730065da40Cited by top-tier papers7
- Do Residual Neural Networks discretize Neural Ordinary Differential Equations?Michael E. Sander, Pierre Ablin, Gabriel PeyréNeurIPS 2022 · 42 citations
- Generalization bounds for neural ordinary differential equations and deep residual networksPierre MarionNeurIPS 2023 · 37 citations
- Neural signature kernels as infinite-width-depth-limits of controlled ResNetsNicola Muca Cirone, Maud Lemercier, Cristopher SalviICML 2023 · 33 citations
- Implicit regularization of deep residual networks towards neural ODEsPierre Marion, Yu-Han Wu, Michael Eli Sander, Gérard BiauICLR 2024 · 24 citations
- Residual Alignment: Uncovering the Mechanisms of Residual NetworksJianing Li, Vardan PapyanNeurIPS 2023 · 21 citations
Builds on2
- A Mean Field Analysis Of Deep ResNet And Beyond: Towards Provably Optimization Via Overparameterization From DepthYiping Lu, Chao Ma, Yulong Lu, Jianfeng Lu et al.ICML 2020 · 85 citations
- ResNet After All: Neural ODEs and Their Numerical SolutionKatharina Ott, Prateek Katiyar, Philipp Hennig, Michael TiemannICLR 2021 · 34 citations
Related papers
- The Neural Covariance SDE: Shaped Infinite Depth-and-Width Networks at InitializationMufan Bill Li, Mihai Nica, Daniel M. RoyNeurIPS 2022 · 51 citations
- Global Convergence in Neural ODEs: Impact of Activation FunctionsTianxiang Gao, Siyuan Sun, Hailiang Liu, Hongyang GaoICLR 2025
- The future is log-Gaussian: ResNets and their infinite-depth-and-width limit at initializationMufan Bill Li, Mihai Nica, Daniel M. RoyNeurIPS 2021 · 41 citations
- On Robustness of Neural Ordinary Differential EquationsHanshu Yan, Jiawei Du, Vincent Y. F. Tan, Jiashi FengICLR 2020 · 161 citations
- Characterizing ResNet's Universal Approximation CapabilityChenghao Liu, Enming Liang, Minghua ChenICML 2024
