Symplectic Adjoint Method for Exact Gradient of Neural ODE with Minimal Memory
Takashi Matsubara, Yuto Miyatake, Takaharu Yaguchi
摘要
A neural network model of a differential equation, namely neural ODE, has enabled the learning of continuous-time dynamical systems and probabilistic distributions with high accuracy. The neural ODE uses the same network repeatedly during a numerical integration. The memory consumption of the backpropagation algorithm is proportional to the number of uses times the network size. This is true even if a checkpointing scheme divides the computation graph into sub-graphs. Otherwise, the adjoint method obtains a gradient by a numerical integration backward in time. Although this method consumes memory only for a single network use, it requires high computational cost to suppress numerical errors. This study proposes the symplectic adjoint method, which is an adjoint method solved by a symplectic integrator. The symplectic adjoint method obtains the exact gradient (up to rounding error) with memory proportional to the number of uses plus the network size. The experimental results demonstrate that the symplectic adjoint method consumes much less memory than the naive backpropagation algorithm and checkpointing schemes, performs faster than the adjoint method, and is more robust to rounding errors.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Learning Neural Event Functions for Ordinary Differential EquationsRicky T. Q. Chen, Brandon Amos, Maximilian NickelICLR 2021 · 被引用 24 次
- Homotopy-based training of NeuralODEs for accurate dynamics discoveryJoon-Hyuk Ko, Hankyul Koh, Nojun Park, Wonho JheNeurIPS 2023 · 被引用 23 次
- Symplectic Spectrum Gaussian Processes: Learning Hamiltonians from Noisy and Sparse DataYusuke Tanaka, Tomoharu Iwata, Naonori UedaNeurIPS 2022 · 被引用 16 次
- Improving Neural ODE Training with Temporal Adaptive Batch NormalizationSu Zheng, Zhengqi Gao, Fan-Keng Sun, Duane S. Boning 等NeurIPS 2024 · 被引用 5 次
- Rex: A Family of Reversible Exponential (Stochastic) Runge-Kutta SolversZander Blasingame, Chen LiuICML 2026 · 被引用 2 次
它引用的顶会 Paper7
- Neural Controlled Differential Equations for Irregular Time SeriesPatrick Kidger, James Morrill, James Foster, Terry J. LyonsNeurIPS 2020 · 被引用 850 次
- PointFlow: 3D Point Cloud Generation With Continuous Normalizing FlowsGuandao Yang, Xun Huang, Zekun Hao, Ming-Yu Liu 等ICCV 2019 · 被引用 794 次
- Symplectic Recurrent Neural NetworksZhengdao Chen, Jianyu Zhang, Martín Arjovsky, Léon BottouICLR 2020 · 被引用 261 次
- Adaptive Checkpoint Adjoint Method for Gradient Estimation in Neural ODEJuntang Zhuang, Nicha C. Dvornek, Xiaoxiao Li, Sekhar Tatikonda 等ICML 2020 · 被引用 125 次
- MALI: A memory efficient and reverse accurate integrator for Neural ODEsJuntang Zhuang, Nicha C. Dvornek, Sekhar Tatikonda, James S. DuncanICLR 2021 · 被引用 61 次
相关 Paper
- Interpolation Technique to Speed Up Gradients Propagation in Neural ODEsTalgat Daulbaev, Alexandr Katrutsa, Larisa Markeeva, Julia Gusak 等NeurIPS 2020 · 被引用 26 次
- Do Residual Neural Networks discretize Neural Ordinary Differential Equations?Michael E. Sander, Pierre Ablin, Gabriel PeyréNeurIPS 2022 · 被引用 42 次
- "Hey, that's not an ODE": Faster ODE Adjoints via SeminormsPatrick Kidger, Ricky T. Q. Chen, Terry J. LyonsICML 2021 · 被引用 56 次
- Efficient Training of Neural Fractional-Order Differential Equation via Adjoint BackpropagationQiyu Kang, Xuhao Li, Kai Zhao, Wenjun Cui 等AAAI 2025 · 被引用 6 次
- Opening the Blackbox: Accelerating Neural Differential Equations by Regularizing Internal Solver HeuristicsAvik Pal, Yingbo Ma, Viral B. Shah, Christopher Vincent RackauckasICML 2021 · 被引用 44 次
