Implicit regularization of deep residual networks towards neural ODEs
Pierre Marion, Yu-Han Wu, Michael Eli Sander, Gérard Biau
Abstract
Residual neural networks are state-of-the-art deep learning models. Their continuous-depth analog, neural ordinary differential equations (ODEs), are also widely used. Despite their success, the link between the discrete and continuous models still lacks a solid mathematical foundation. In this article, we take a step in this direction by establishing an implicit regularization of deep residual networks towards neural ODEs, for nonlinear networks trained with gradient flow. We prove that if the network is initialized as a discretization of a neural ODE, then such a discretization holds throughout training. Our results are valid for a finite training time, and also as the training time tends to infinity provided that the network satisfies a Polyak-Lojasiewicz condition. Importantly, this condition holds for a family of residual networks where the residuals are two-layer perceptrons with an overparameterization in width that is only linear, and implies the convergence of gradient flow to a global minimum. Numerical experiments illustrate our results.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers13
- Generalization bounds for neural ordinary differential equations and deep residual networksPierre MarionNeurIPS 2023 · 37 citations
- Deep linear networks for regression are implicitly regularized towards flat minimaPierre Marion, Lénaïc ChizatNeurIPS 2024 · 21 citations
- EGODE: An Event-attended Graph ODE Framework for Modeling Rigid DynamicsJingyang Yuan, Gongbo Sun, Zhiping Xiao, Hang Zhou et al.NeurIPS 2024 · 11 citations
- Clustering in Deep Stochastic TransformersLev Fedorov, Michael Sander, Romuald Elie, Pierre Marion et al.ICML 2026 · 7 citations
- Do Neural Networks Need Gradient Descent to Generalize? A Theoretical StudyYotam Alexander, Yonatan Slutzky, Yuval Ran-Milo, Nadav CohenNeurIPS 2025 · 3 citations
Builds on17
- Gradient Descent Maximizes the Margin of Homogeneous Neural NetworksKaifeng Lyu, Jian LiICLR 2020 · 402 citations
- Dissecting Neural ODEsStefano Massaroli, Michael Poli, Jinkyoo Park, Atsushi Yamashita et al.NeurIPS 2020 · 261 citations
- The Lipschitz Constant of Self-AttentionHyunjik Kim, George Papamakarios, Andriy MnihICML 2021 · 208 citations
- Batch Normalization Biases Residual Blocks Towards the Identity Function in Deep NetworksSoham De, Samuel L. SmithNeurIPS 2020 · 173 citations
- The emergence of clusters in self-attention dynamicsBorjan Geshkovski, Cyril Letrouit, Yury Polyanskiy, Philippe RigolletNeurIPS 2023 · 163 citations
Related papers
- Do Residual Neural Networks discretize Neural Ordinary Differential Equations?Michael E. Sander, Pierre Ablin, Gabriel PeyréNeurIPS 2022 · 42 citations
- Scaling Properties of Deep Residual NetworksAlain-Sam Cohen, Rama Cont, Alain Rossier, Renyuan XuICML 2021 · 21 citations
- A global convergence theory for deep ReLU implicit networks via over-parameterizationTianxiang Gao, Hailiang Liu, Jia Liu, Hridesh Rajan et al.ICLR 2022 · 21 citations
- Imbedding Deep Neural NetworksAndrew Corbett, Dmitry KanginICLR 2022 · 2 citations
- On Robustness of Neural Ordinary Differential EquationsHanshu Yan, Jiawei Du, Vincent Y. F. Tan, Jiashi FengICLR 2020 · 161 citations
