MALI: A memory efficient and reverse accurate integrator for Neural ODEs
Juntang Zhuang, Nicha C. Dvornek, Sekhar Tatikonda, James S. Duncan
摘要
Neural ordinary differential equations (Neural ODEs) are a new family of deep-learning models with continuous depth. However, the numerical estimation of the gradient in the continuous case is not well solved: existing implementations of the adjoint method suffer from inaccuracy in reverse-time trajectory, while the naive method and the adaptive checkpoint adjoint method (ACA) have a memory cost that grows with integration time. In this project, based on the asynchronous leapfrog (ALF) solver, we propose the Memory-efficient ALF Integrator (MALI), which has a constant memory cost w.r.t number of solver steps in integration similar to the adjoint method, and guarantees accuracy in reverse-time trajectory (hence accuracy in gradient estimation). We validate MALI in various tasks: on image recognition tasks, to our knowledge, MALI is the first to enable feasible training of a Neural ODE on ImageNet and outperform a well-tuned ResNet, while existing methods fail due to either heavy memory burden or inaccuracy; for time series modeling, MALI significantly outperforms the adjoint method; and for continuous generative models, MALI achieves new state-of-the-art performance. We provide a pypi package at https://jzkay12.github.io/TorchDiffEqPack/
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Efficient and Accurate Gradients for Neural SDEsPatrick Kidger, James Foster, Xuechen Li, Terry J. LyonsNeurIPS 2021 · 被引用 107 次
- Heavy Ball Neural Ordinary Differential EquationsHedi Xia, Vai Suliafu, Hangjie Ji, Tan M. Nguyen 等NeurIPS 2021 · 被引用 75 次
- Opening the Blackbox: Accelerating Neural Differential Equations by Regularizing Internal Solver HeuristicsAvik Pal, Yingbo Ma, Viral B. Shah, Christopher Vincent RackauckasICML 2021 · 被引用 44 次
- EXIT: Extrapolation and Interpolation-based Neural Controlled Differential Equations for Time-series Classification and ForecastingSheo Yon Jhin, Jaehoon Lee, Minju Jo, Seungji Kook 等WWW 2022 · 被引用 30 次
- Symplectic Adjoint Method for Exact Gradient of Neural ODE with Minimal MemoryTakashi Matsubara, Yuto Miyatake, Takaharu YaguchiNeurIPS 2021 · 被引用 30 次
它引用的顶会 Paper5
- Neural Controlled Differential Equations for Irregular Time SeriesPatrick Kidger, James Morrill, James Foster, Terry J. LyonsNeurIPS 2020 · 被引用 850 次
- Dissecting Neural ODEsStefano Massaroli, Michael Poli, Jinkyoo Park, Atsushi Yamashita 等NeurIPS 2020 · 被引用 261 次
- On Robustness of Neural Ordinary Differential EquationsHanshu Yan, Jiawei Du, Vincent Y. F. Tan, Jiashi FengICLR 2020 · 被引用 161 次
- Adaptive Checkpoint Adjoint Method for Gradient Estimation in Neural ODEJuntang Zhuang, Nicha C. Dvornek, Xiaoxiao Li, Sekhar Tatikonda 等ICML 2020 · 被引用 125 次
- How to Train Your Neural ODE: the World of Jacobian and Kinetic RegularizationChris Finlay, Jörn-Henrik Jacobsen, Levon Nurbekyan, Adam M. ObermanICML 2020 · 被引用 76 次
相关 Paper
- Do Residual Neural Networks discretize Neural Ordinary Differential Equations?Michael E. Sander, Pierre Ablin, Gabriel PeyréNeurIPS 2022 · 被引用 42 次
- Second-Order Neural ODE OptimizerGuan-Horng Liu, Tianrong Chen, Evangelos A. TheodorouNeurIPS 2021 · 被引用 20 次
- Locally Regularized Neural Differential Equations: Some Black Boxes were meant to remain closed!Avik Pal, Alan Edelman, Christopher Vincent RackauckasICML 2023 · 被引用 4 次
- Interpolation Technique to Speed Up Gradients Propagation in Neural ODEsTalgat Daulbaev, Alexandr Katrutsa, Larisa Markeeva, Julia Gusak 等NeurIPS 2020 · 被引用 26 次
- AdjointDEIS: Efficient Gradients for Diffusion ModelsZander W. Blasingame, Chen LiuNeurIPS 2024 · 被引用 8 次
