Do Residual Neural Networks discretize Neural Ordinary Differential Equations?
Michael E. Sander, Pierre Ablin, Gabriel Peyré
Abstract
Neural Ordinary Differential Equations (Neural ODEs) are the continuous analog of Residual Neural Networks (ResNets). We investigate whether the discrete dynamics defined by a ResNet are close to the continuous one of a Neural ODE. We first quantify the distance between the ResNet's hidden state trajectory and the solution of its corresponding Neural ODE. Our bound is tight and, on the negative side, does not go to 0 with depth N if the residual functions are not smooth with depth. On the positive side, we show that this smoothness is preserved by gradient descent for a ResNet with linear residual functions and small enough initial loss. It ensures an implicit regularization towards a limit Neural ODE at rate 1 over N, uniformly with depth and optimization time. As a byproduct of our analysis, we consider the use of a memory-free discrete adjoint method to train a ResNet by recovering the activations on the fly through a backward pass of the network, and show that this method theoretically succeeds at large depth if the residual functions are Lipschitz with the input. We then show that Heun's method, a second order ODE integration scheme, allows for better gradient estimation with the adjoint method when the residual functions are smooth with depth. We experimentally validate that our adjoint method succeeds at large depth, and that Heun method needs fewer layers to succeed. We finally use the adjoint method successfully for fine-tuning very deep ResNets without memory consumption in the residual layers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d3b6df1c-c384-4aef-bfa4-8c95984b4e29Cited by top-tier papers12
- Generalization bounds for neural ordinary differential equations and deep residual networksPierre MarionNeurIPS 2023 · 37 citations
- Implicit regularization of deep residual networks towards neural ODEsPierre Marion, Yu-Han Wu, Michael Eli Sander, Gérard BiauICLR 2024 · 24 citations
- Deep linear networks for regression are implicitly regularized towards flat minimaPierre Marion, Lénaïc ChizatNeurIPS 2024 · 21 citations
- Topological Neural Networks go Persistent, Equivariant, and ContinuousYogesh Verma, Amauri H. Souza, Vikas GargICML 2024 · 13 citations
- Vocabulary for Universal Approximation: A Linguistic Perspective of Mapping CompositionsYongqiang CaiICML 2024 · 11 citations
Builds on8
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Revisiting ResNets: Improved Training and Scaling StrategiesIrwan Bello, William Fedus, Xianzhi Du, Ekin Dogus Cubuk et al.NeurIPS 2021 · 378 citations
- On the linearity of large non-linear models: when and why the tangent kernel is constantChaoyue Liu, Libin Zhu, Mikhail BelkinNeurIPS 2020 · 183 citations
- Adaptive Checkpoint Adjoint Method for Gradient Estimation in Neural ODEJuntang Zhuang, Nicha C. Dvornek, Xiaoxiao Li, Sekhar Tatikonda et al.ICML 2020 · 125 citations
- A Mean Field Analysis Of Deep ResNet And Beyond: Towards Provably Optimization Via Overparameterization From DepthYiping Lu, Chao Ma, Yulong Lu, Jianfeng Lu et al.ICML 2020 · 85 citations
Related papers
- Scaling Properties of Deep Residual NetworksAlain-Sam Cohen, Rama Cont, Alain Rossier, Renyuan XuICML 2021 · 21 citations
- MALI: A memory efficient and reverse accurate integrator for Neural ODEsJuntang Zhuang, Nicha C. Dvornek, Sekhar Tatikonda, James S. DuncanICLR 2021 · 61 citations
- Symplectic Adjoint Method for Exact Gradient of Neural ODE with Minimal MemoryTakashi Matsubara, Yuto Miyatake, Takaharu YaguchiNeurIPS 2021 · 30 citations
- Imbedding Deep Neural NetworksAndrew Corbett, Dmitry KanginICLR 2022 · 2 citations
- Interpolation Technique to Speed Up Gradients Propagation in Neural ODEsTalgat Daulbaev, Alexandr Katrutsa, Larisa Markeeva, Julia Gusak et al.NeurIPS 2020 · 26 citations
