Improving Neural Ordinary Differential Equations with Nesterov's Accelerated Gradient Method
Ho Huu Nghia Nguyen, Tan Nguyen, Huyen Vo, Stanley J. Osher, Thieu Vo
Abstract
We propose the Nesterov neural ordinary differential equations (NesterovNODEs), whose layers solve the second-order ordinary differential equations (ODEs) limit of Nesterov’s accelerated gradient (NAG) method, and a generalization called GNesterovNODEs. Taking the advantage of the convergence rate O (1 /k 2 ) of the NAG scheme, GNesterovNODEs speed up training and inference by reducing the number of function evaluations (NFEs) needed to solve the ODEs. We also prove that the adjoint state of a GNesterovNODEs also satisfies a GNesterovNODEs, thus accelerating both forward and backward ODE solvers and allowing the model to be scaled up for large-scale tasks. We empirically corroborate the advantage of GNesterovNODEs on a wide range of practical applications, including point cloud separation, image classification, and sequence modeling. Compared to NODEs, GNesterovNODEs require a significantly smaller number of NFEs while achieving better accuracy across our experiments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5f5beb44-b8e4-4082-b097-eb97b8cf1e91Cited by top-tier papers7
- MomentumSMoE: Integrating Momentum into Sparse Mixture of ExpertsRachel S. Y. Teo, Tan M. NguyenNeurIPS 2024 · 10 citations
- Expert Merging in Sparse Mixture of Experts with Nash BargainingDung Viet Nguyen, Anh Nguyen Thi, Minh Hoang Nguyen, Luc Nguyen et al.ICLR 2026 · 3 citations
- Operator-Learning-Inspired Modeling of Neural Ordinary Differential EquationsWoojin Cho, Seunghyeon Cho, Hyundong Jin, Jinsung Jeon et al.AAAI 2024 · 3 citations
- Unsupervised Trajectory Optimization for 3D Registration in Serial Section Electron Microscopy using Neural ODEsZhenbang Zhang, Jingtong Feng, Hongjia Li, Haythem El-Messiry et al.NeurIPS 2025 · 1 citation
- Demystifying the Token Dynamics of Deep Selective State Space ModelsThieu N. Vo, Duy-Tung Pham, Xin T. Tong, Tan Minh NguyenICLR 2025
Builds on13
- Equivariant Flows: Exact Likelihood Generative Learning for Symmetric DensitiesJonas Köhler, Leon Klein, Frank NoéICML 2020 · 330 citations
- Symplectic ODE-Net: Learning Hamiltonian Dynamics with ControlYaofeng Desmond Zhong, Biswadip Dey, Amit ChakrabortyICLR 2020 · 319 citations
- Dissecting Neural ODEsStefano Massaroli, Michael Poli, Jinkyoo Park, Atsushi Yamashita et al.NeurIPS 2020 · 261 citations
- Learning Differential Equations that are Easy to SolveJacob Kelly, Jesse Bettencourt, Matthew J. Johnson, David DuvenaudNeurIPS 2020 · 134 citations
- On Second Order Behaviour in Augmented Neural ODEsAlexander Norcliffe, Cristian Bodnar, Ben Day, Nikola Simidjievski et al.NeurIPS 2020 · 116 citations
Related papers
- Heavy Ball Neural Ordinary Differential EquationsHedi Xia, Vai Suliafu, Hangjie Ji, Tan M. Nguyen et al.NeurIPS 2021 · 75 citations
- Adaptive Checkpoint Adjoint Method for Gradient Estimation in Neural ODEJuntang Zhuang, Nicha C. Dvornek, Xiaoxiao Li, Sekhar Tatikonda et al.ICML 2020 · 125 citations
- Second-Order Neural ODE OptimizerGuan-Horng Liu, Tianrong Chen, Evangelos A. TheodorouNeurIPS 2021 · 20 citations
- Neural Delay Differential EquationsQunxi Zhu, Yao Guo, Wei LinICLR 2021 · 3 citations
- Do Residual Neural Networks discretize Neural Ordinary Differential Equations?Michael E. Sander, Pierre Ablin, Gabriel PeyréNeurIPS 2022 · 42 citations
