Improving Neural Ordinary Differential Equations with Nesterov's Accelerated Gradient Method
Ho Huu Nghia Nguyen, Tan Nguyen, Huyen Vo, Stanley J. Osher, Thieu Vo
摘要
We propose the Nesterov neural ordinary differential equations (NesterovNODEs), whose layers solve the second-order ordinary differential equations (ODEs) limit of Nesterov’s accelerated gradient (NAG) method, and a generalization called GNesterovNODEs. Taking the advantage of the convergence rate O (1 /k 2 ) of the NAG scheme, GNesterovNODEs speed up training and inference by reducing the number of function evaluations (NFEs) needed to solve the ODEs. We also prove that the adjoint state of a GNesterovNODEs also satisfies a GNesterovNODEs, thus accelerating both forward and backward ODE solvers and allowing the model to be scaled up for large-scale tasks. We empirically corroborate the advantage of GNesterovNODEs on a wide range of practical applications, including point cloud separation, image classification, and sequence modeling. Compared to NODEs, GNesterovNODEs require a significantly smaller number of NFEs while achieving better accuracy across our experiments.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- MomentumSMoE: Integrating Momentum into Sparse Mixture of ExpertsRachel S. Y. Teo, Tan M. NguyenNeurIPS 2024 · 被引用 10 次
- Expert Merging in Sparse Mixture of Experts with Nash BargainingDung Viet Nguyen, Anh Nguyen Thi, Minh Hoang Nguyen, Luc Nguyen 等ICLR 2026 · 被引用 3 次
- Operator-Learning-Inspired Modeling of Neural Ordinary Differential EquationsWoojin Cho, Seunghyeon Cho, Hyundong Jin, Jinsung Jeon 等AAAI 2024 · 被引用 3 次
- Unsupervised Trajectory Optimization for 3D Registration in Serial Section Electron Microscopy using Neural ODEsZhenbang Zhang, Jingtong Feng, Hongjia Li, Haythem El-Messiry 等NeurIPS 2025 · 被引用 1 次
- Demystifying the Token Dynamics of Deep Selective State Space ModelsThieu N. Vo, Duy-Tung Pham, Xin T. Tong, Tan Minh NguyenICLR 2025
它引用的顶会 Paper13
- Equivariant Flows: Exact Likelihood Generative Learning for Symmetric DensitiesJonas Köhler, Leon Klein, Frank NoéICML 2020 · 被引用 330 次
- Symplectic ODE-Net: Learning Hamiltonian Dynamics with ControlYaofeng Desmond Zhong, Biswadip Dey, Amit ChakrabortyICLR 2020 · 被引用 319 次
- Dissecting Neural ODEsStefano Massaroli, Michael Poli, Jinkyoo Park, Atsushi Yamashita 等NeurIPS 2020 · 被引用 261 次
- Learning Differential Equations that are Easy to SolveJacob Kelly, Jesse Bettencourt, Matthew J. Johnson, David DuvenaudNeurIPS 2020 · 被引用 134 次
- On Second Order Behaviour in Augmented Neural ODEsAlexander Norcliffe, Cristian Bodnar, Ben Day, Nikola Simidjievski 等NeurIPS 2020 · 被引用 116 次
相关 Paper
- Heavy Ball Neural Ordinary Differential EquationsHedi Xia, Vai Suliafu, Hangjie Ji, Tan M. Nguyen 等NeurIPS 2021 · 被引用 75 次
- Adaptive Checkpoint Adjoint Method for Gradient Estimation in Neural ODEJuntang Zhuang, Nicha C. Dvornek, Xiaoxiao Li, Sekhar Tatikonda 等ICML 2020 · 被引用 125 次
- Second-Order Neural ODE OptimizerGuan-Horng Liu, Tianrong Chen, Evangelos A. TheodorouNeurIPS 2021 · 被引用 20 次
- Neural Delay Differential EquationsQunxi Zhu, Yao Guo, Wei LinICLR 2021 · 被引用 3 次
- Do Residual Neural Networks discretize Neural Ordinary Differential Equations?Michael E. Sander, Pierre Ablin, Gabriel PeyréNeurIPS 2022 · 被引用 42 次
