Second-Order Neural ODE Optimizer
Guan-Horng Liu, Tianrong Chen, Evangelos A. Theodorou
Abstract
We propose a novel second-order optimization framework for training the emerging deep continuous-time models, specifically the Neural Ordinary Differential Equations (Neural ODEs). Since their training already involves expensive gradient computation by solving a backward ODE, deriving efficient second-order methods becomes highly nontrivial. Nevertheless, inspired by the recent Optimal Control (OC) interpretation of training deep networks, we show that a specific continuous-time OC methodology, called Differential Programming, can be adopted to derive backward ODEs for higher-order derivatives at the same O(1) memory cost. We further explore a low-rank representation of the second-order derivatives and show that it leads to efficient preconditioned updates with the aid of Kronecker-based factorization. The resulting method -- named SNOpt -- converges much faster than first-order baselines in wall-clock time, and the improvement remains consistent across various applications, e.g. image classification, generative flow, and time-series prediction. Our framework also enables direct architecture optimization, such as the integration time of Neural ODEs, with second-order feedback policies, strengthening the OC perspective as a principled tool of analyzing optimization in deep learning. Our code is available at https://github.com/ghliu/snopt.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4e57e7da-6f80-45cd-be11-f3b69c128a73Cited by top-tier papers5
- Deep Generalized Schrödinger BridgeGuan-Horng Liu, Tianrong Chen, Oswin So, Evangelos A. TheodorouNeurIPS 2022 · 64 citations
- Flexible Language Modeling in Continuous Space with Transformer-based Autoregressive FlowsRuixiang Zhang, Shuangfei Zhai, Jiatao Gu, Yizhe Zhang et al.NeurIPS 2025 · 8 citations
- How Deep Do We Need: Accelerating Training and Inference of Neural ODEs via Control PerspectiveKeyan Miao, Konstantinos GatsisICML 2024 · 2 citations
- Normalizing Flows are Capable Generative ModelsShuangfei Zhai, Ruixiang Zhang, Preetum Nakkiran, David Berthelot et al.ICML 2025
- A robust differential Neural ODE OptimizerPanagiotis Theodoropoulos, Guan-Horng Liu, Tianrong Chen, Augustinos D. Saravanos et al.ICLR 2024
Builds on13
- Neural Controlled Differential Equations for Irregular Time SeriesPatrick Kidger, James Morrill, James Foster, Terry J. LyonsNeurIPS 2020 · 850 citations
- Symplectic ODE-Net: Learning Hamiltonian Dynamics with ControlYaofeng Desmond Zhong, Biswadip Dey, Amit ChakrabortyICLR 2020 · 319 citations
- Dissecting Neural ODEsStefano Massaroli, Michael Poli, Jinkyoo Park, Atsushi Yamashita et al.NeurIPS 2020 · 261 citations
- OT-Flow: Fast and Accurate Continuous Normalizing Flows via Optimal TransportDerek Onken, Samy Wu Fung, Xingjian Li, Lars RuthottoAAAI 2021 · 210 citations
- Riemannian Continuous Normalizing FlowsEmile Mathieu, Maximilian NickelNeurIPS 2020 · 198 citations
Related papers
- DDPNOpt: Differential Dynamic Programming Neural OptimizerGuan-Horng Liu, Tianrong Chen, Evangelos A. TheodorouICLR 2021 · 7 citations
- Locally Regularized Neural Differential Equations: Some Black Boxes were meant to remain closed!Avik Pal, Alan Edelman, Christopher Vincent RackauckasICML 2023 · 4 citations
- MALI: A memory efficient and reverse accurate integrator for Neural ODEsJuntang Zhuang, Nicha C. Dvornek, Sekhar Tatikonda, James S. DuncanICLR 2021 · 61 citations
- HNO: High-Order Numerical Architecture for ODE-Inspired Deep Unfolding NetworksLin Kong, Wei Sun, Fanhua Shang, Yuanyuan Liu et al.AAAI 2022 · 1 citation
- Improving Neural Ordinary Differential Equations with Nesterov's Accelerated Gradient MethodHo Huu Nghia Nguyen, Tan Nguyen, Huyen Vo, Stanley J. Osher et al.NeurIPS 2022 · 28 citations
