Do Residual Neural Networks discretize Neural Ordinary Differential Equations?
Michael E. Sander, Pierre Ablin, Gabriel Peyré
摘要
Neural Ordinary Differential Equations (Neural ODEs) are the continuous analog of Residual Neural Networks (ResNets). We investigate whether the discrete dynamics defined by a ResNet are close to the continuous one of a Neural ODE. We first quantify the distance between the ResNet's hidden state trajectory and the solution of its corresponding Neural ODE. Our bound is tight and, on the negative side, does not go to 0 with depth N if the residual functions are not smooth with depth. On the positive side, we show that this smoothness is preserved by gradient descent for a ResNet with linear residual functions and small enough initial loss. It ensures an implicit regularization towards a limit Neural ODE at rate 1 over N, uniformly with depth and optimization time. As a byproduct of our analysis, we consider the use of a memory-free discrete adjoint method to train a ResNet by recovering the activations on the fly through a backward pass of the network, and show that this method theoretically succeeds at large depth if the residual functions are Lipschitz with the input. We then show that Heun's method, a second order ODE integration scheme, allows for better gradient estimation with the adjoint method when the residual functions are smooth with depth. We experimentally validate that our adjoint method succeeds at large depth, and that Heun method needs fewer layers to succeed. We finally use the adjoint method successfully for fine-tuning very deep ResNets without memory consumption in the residual layers.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Generalization bounds for neural ordinary differential equations and deep residual networksPierre MarionNeurIPS 2023 · 被引用 37 次
- Implicit regularization of deep residual networks towards neural ODEsPierre Marion, Yu-Han Wu, Michael Eli Sander, Gérard BiauICLR 2024 · 被引用 24 次
- Deep linear networks for regression are implicitly regularized towards flat minimaPierre Marion, Lénaïc ChizatNeurIPS 2024 · 被引用 21 次
- Topological Neural Networks go Persistent, Equivariant, and ContinuousYogesh Verma, Amauri H. Souza, Vikas GargICML 2024 · 被引用 13 次
- Vocabulary for Universal Approximation: A Linguistic Perspective of Mapping CompositionsYongqiang CaiICML 2024 · 被引用 11 次
它引用的顶会 Paper8
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Revisiting ResNets: Improved Training and Scaling StrategiesIrwan Bello, William Fedus, Xianzhi Du, Ekin Dogus Cubuk 等NeurIPS 2021 · 被引用 378 次
- On the linearity of large non-linear models: when and why the tangent kernel is constantChaoyue Liu, Libin Zhu, Mikhail BelkinNeurIPS 2020 · 被引用 183 次
- Adaptive Checkpoint Adjoint Method for Gradient Estimation in Neural ODEJuntang Zhuang, Nicha C. Dvornek, Xiaoxiao Li, Sekhar Tatikonda 等ICML 2020 · 被引用 125 次
- A Mean Field Analysis Of Deep ResNet And Beyond: Towards Provably Optimization Via Overparameterization From DepthYiping Lu, Chao Ma, Yulong Lu, Jianfeng Lu 等ICML 2020 · 被引用 85 次
相关 Paper
- Scaling Properties of Deep Residual NetworksAlain-Sam Cohen, Rama Cont, Alain Rossier, Renyuan XuICML 2021 · 被引用 21 次
- MALI: A memory efficient and reverse accurate integrator for Neural ODEsJuntang Zhuang, Nicha C. Dvornek, Sekhar Tatikonda, James S. DuncanICLR 2021 · 被引用 61 次
- Symplectic Adjoint Method for Exact Gradient of Neural ODE with Minimal MemoryTakashi Matsubara, Yuto Miyatake, Takaharu YaguchiNeurIPS 2021 · 被引用 30 次
- Imbedding Deep Neural NetworksAndrew Corbett, Dmitry KanginICLR 2022 · 被引用 2 次
- Interpolation Technique to Speed Up Gradients Propagation in Neural ODEsTalgat Daulbaev, Alexandr Katrutsa, Larisa Markeeva, Julia Gusak 等NeurIPS 2020 · 被引用 26 次
