Second-Order Neural ODE Optimizer
Guan-Horng Liu, Tianrong Chen, Evangelos A. Theodorou
摘要
We propose a novel second-order optimization framework for training the emerging deep continuous-time models, specifically the Neural Ordinary Differential Equations (Neural ODEs). Since their training already involves expensive gradient computation by solving a backward ODE, deriving efficient second-order methods becomes highly nontrivial. Nevertheless, inspired by the recent Optimal Control (OC) interpretation of training deep networks, we show that a specific continuous-time OC methodology, called Differential Programming, can be adopted to derive backward ODEs for higher-order derivatives at the same O(1) memory cost. We further explore a low-rank representation of the second-order derivatives and show that it leads to efficient preconditioned updates with the aid of Kronecker-based factorization. The resulting method -- named SNOpt -- converges much faster than first-order baselines in wall-clock time, and the improvement remains consistent across various applications, e.g. image classification, generative flow, and time-series prediction. Our framework also enables direct architecture optimization, such as the integration time of Neural ODEs, with second-order feedback policies, strengthening the OC perspective as a principled tool of analyzing optimization in deep learning. Our code is available at https://github.com/ghliu/snopt.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Deep Generalized Schrödinger BridgeGuan-Horng Liu, Tianrong Chen, Oswin So, Evangelos A. TheodorouNeurIPS 2022 · 被引用 64 次
- Flexible Language Modeling in Continuous Space with Transformer-based Autoregressive FlowsRuixiang Zhang, Shuangfei Zhai, Jiatao Gu, Yizhe Zhang 等NeurIPS 2025 · 被引用 8 次
- How Deep Do We Need: Accelerating Training and Inference of Neural ODEs via Control PerspectiveKeyan Miao, Konstantinos GatsisICML 2024 · 被引用 2 次
- Normalizing Flows are Capable Generative ModelsShuangfei Zhai, Ruixiang Zhang, Preetum Nakkiran, David Berthelot 等ICML 2025
- A robust differential Neural ODE OptimizerPanagiotis Theodoropoulos, Guan-Horng Liu, Tianrong Chen, Augustinos D. Saravanos 等ICLR 2024
它引用的顶会 Paper13
- Neural Controlled Differential Equations for Irregular Time SeriesPatrick Kidger, James Morrill, James Foster, Terry J. LyonsNeurIPS 2020 · 被引用 850 次
- Symplectic ODE-Net: Learning Hamiltonian Dynamics with ControlYaofeng Desmond Zhong, Biswadip Dey, Amit ChakrabortyICLR 2020 · 被引用 319 次
- Dissecting Neural ODEsStefano Massaroli, Michael Poli, Jinkyoo Park, Atsushi Yamashita 等NeurIPS 2020 · 被引用 261 次
- OT-Flow: Fast and Accurate Continuous Normalizing Flows via Optimal TransportDerek Onken, Samy Wu Fung, Xingjian Li, Lars RuthottoAAAI 2021 · 被引用 210 次
- Riemannian Continuous Normalizing FlowsEmile Mathieu, Maximilian NickelNeurIPS 2020 · 被引用 198 次
相关 Paper
- DDPNOpt: Differential Dynamic Programming Neural OptimizerGuan-Horng Liu, Tianrong Chen, Evangelos A. TheodorouICLR 2021 · 被引用 7 次
- Locally Regularized Neural Differential Equations: Some Black Boxes were meant to remain closed!Avik Pal, Alan Edelman, Christopher Vincent RackauckasICML 2023 · 被引用 4 次
- MALI: A memory efficient and reverse accurate integrator for Neural ODEsJuntang Zhuang, Nicha C. Dvornek, Sekhar Tatikonda, James S. DuncanICLR 2021 · 被引用 61 次
- HNO: High-Order Numerical Architecture for ODE-Inspired Deep Unfolding NetworksLin Kong, Wei Sun, Fanhua Shang, Yuanyuan Liu 等AAAI 2022 · 被引用 1 次
- Improving Neural Ordinary Differential Equations with Nesterov's Accelerated Gradient MethodHo Huu Nghia Nguyen, Tan Nguyen, Huyen Vo, Stanley J. Osher 等NeurIPS 2022 · 被引用 28 次
