Symplectic Adjoint Method for Exact Gradient of Neural ODE with Minimal Memory
Takashi Matsubara, Yuto Miyatake, Takaharu Yaguchi
Abstract
A neural network model of a differential equation, namely neural ODE, has enabled the learning of continuous-time dynamical systems and probabilistic distributions with high accuracy. The neural ODE uses the same network repeatedly during a numerical integration. The memory consumption of the backpropagation algorithm is proportional to the number of uses times the network size. This is true even if a checkpointing scheme divides the computation graph into sub-graphs. Otherwise, the adjoint method obtains a gradient by a numerical integration backward in time. Although this method consumes memory only for a single network use, it requires high computational cost to suppress numerical errors. This study proposes the symplectic adjoint method, which is an adjoint method solved by a symplectic integrator. The symplectic adjoint method obtains the exact gradient (up to rounding error) with memory proportional to the number of uses plus the network size. The experimental results demonstrate that the symplectic adjoint method consumes much less memory than the naive backpropagation algorithm and checkpointing schemes, performs faster than the adjoint method, and is more robust to rounding errors.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bf5ccc71-5401-4805-8cf6-eacfc19a5961Cited by top-tier papers7
- Learning Neural Event Functions for Ordinary Differential EquationsRicky T. Q. Chen, Brandon Amos, Maximilian NickelICLR 2021 · 24 citations
- Homotopy-based training of NeuralODEs for accurate dynamics discoveryJoon-Hyuk Ko, Hankyul Koh, Nojun Park, Wonho JheNeurIPS 2023 · 23 citations
- Symplectic Spectrum Gaussian Processes: Learning Hamiltonians from Noisy and Sparse DataYusuke Tanaka, Tomoharu Iwata, Naonori UedaNeurIPS 2022 · 16 citations
- Improving Neural ODE Training with Temporal Adaptive Batch NormalizationSu Zheng, Zhengqi Gao, Fan-Keng Sun, Duane S. Boning et al.NeurIPS 2024 · 5 citations
- Rex: A Family of Reversible Exponential (Stochastic) Runge-Kutta SolversZander Blasingame, Chen LiuICML 2026 · 2 citations
Builds on7
- Neural Controlled Differential Equations for Irregular Time SeriesPatrick Kidger, James Morrill, James Foster, Terry J. LyonsNeurIPS 2020 · 850 citations
- PointFlow: 3D Point Cloud Generation With Continuous Normalizing FlowsGuandao Yang, Xun Huang, Zekun Hao, Ming-Yu Liu et al.ICCV 2019 · 794 citations
- Symplectic Recurrent Neural NetworksZhengdao Chen, Jianyu Zhang, Martín Arjovsky, Léon BottouICLR 2020 · 261 citations
- Adaptive Checkpoint Adjoint Method for Gradient Estimation in Neural ODEJuntang Zhuang, Nicha C. Dvornek, Xiaoxiao Li, Sekhar Tatikonda et al.ICML 2020 · 125 citations
- MALI: A memory efficient and reverse accurate integrator for Neural ODEsJuntang Zhuang, Nicha C. Dvornek, Sekhar Tatikonda, James S. DuncanICLR 2021 · 61 citations
Related papers
- Interpolation Technique to Speed Up Gradients Propagation in Neural ODEsTalgat Daulbaev, Alexandr Katrutsa, Larisa Markeeva, Julia Gusak et al.NeurIPS 2020 · 26 citations
- Do Residual Neural Networks discretize Neural Ordinary Differential Equations?Michael E. Sander, Pierre Ablin, Gabriel PeyréNeurIPS 2022 · 42 citations
- "Hey, that's not an ODE": Faster ODE Adjoints via SeminormsPatrick Kidger, Ricky T. Q. Chen, Terry J. LyonsICML 2021 · 56 citations
- Efficient Training of Neural Fractional-Order Differential Equation via Adjoint BackpropagationQiyu Kang, Xuhao Li, Kai Zhao, Wenjun Cui et al.AAAI 2025 · 6 citations
- Opening the Blackbox: Accelerating Neural Differential Equations by Regularizing Internal Solver HeuristicsAvik Pal, Yingbo Ma, Viral B. Shah, Christopher Vincent RackauckasICML 2021 · 44 citations
