Adaptive Checkpoint Adjoint Method for Gradient Estimation in Neural ODE
Juntang Zhuang, Nicha C. Dvornek, Xiaoxiao Li, Sekhar Tatikonda, Xenophon Papademetris, James S. Duncan
Abstract
Neural ordinary differential equations (NODEs) have recently attracted increasing attention; however, their empirical performance on benchmark tasks (e.g. image classification) are significantly inferior to discrete-layer models. We demonstrate an explanation for their poorer performance is the inaccuracy of existing gradient estimation methods: the adjoint method has numerical errors in reverse-mode integration; the naive method directly back-propagates through ODE solvers, but suffers from a redundantly deep computation graph when searching for the optimal stepsize. We propose the Adaptive Checkpoint Adjoint (ACA) method: in automatic differentiation, ACA applies a trajectory checkpoint strategy which records the forward-mode trajectory as the reverse-mode trajectory to guarantee accuracy; ACA deletes redundant components for shallow computation graphs; and ACA supports adaptive solvers. On image classification tasks, compared with the adjoint and naive method, ACA achieves half the error rate in half the training time; NODE trained with ACA outperforms ResNet in both accuracy and test-retest reliability. On time-series modeling, ACA outperforms competing methods. Finally, in an example of the three-body problem, we show NODE with ACA can incorporate physical knowledge to achieve better accuracy. We provide the PyTorch implementation of ACA: https://github.com/ juntang-zhuang/torch-ACA .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c76ebd77-dc62-4088-9a84-ad0bad108b07Cited by top-tier papers41
- Liquid Time-constant NetworksRamin M. Hasani, Mathias Lechner, Alexander Amini, Daniela Rus et al.AAAI 2021 · 399 citations
- Neural Flows: Efficient Alternative to Neural ODEsMarin Bilos, Johanna Sommer, Syama Sundar Rangapuram, Tim Januschowski et al.NeurIPS 2021 · 151 citations
- SE(3) Equivariant Graph Neural Networks with Complete Local FramesWeitao Du, He Zhang, Yuanqi Du, Qi Meng et al.ICML 2022 · 111 citations
- Robust Evaluation of Diffusion-Based Adversarial PurificationMinjong Lee, Dongwoo KimICCV 2023 · 96 citations
- Machine learning structure preserving brackets for forecasting irreversible processesKookjin Lee, Nathaniel Trask, Panos StinisNeurIPS 2021 · 80 citations
Related papers
- MALI: A memory efficient and reverse accurate integrator for Neural ODEsJuntang Zhuang, Nicha C. Dvornek, Sekhar Tatikonda, James S. DuncanICLR 2021 · 61 citations
- Improving Neural Ordinary Differential Equations with Nesterov's Accelerated Gradient MethodHo Huu Nghia Nguyen, Tan Nguyen, Huyen Vo, Stanley J. Osher et al.NeurIPS 2022 · 28 citations
- Symplectic Adjoint Method for Exact Gradient of Neural ODE with Minimal MemoryTakashi Matsubara, Yuto Miyatake, Takaharu YaguchiNeurIPS 2021 · 30 citations
- Heavy Ball Neural Ordinary Differential EquationsHedi Xia, Vai Suliafu, Hangjie Ji, Tan M. Nguyen et al.NeurIPS 2021 · 75 citations
- Do Residual Neural Networks discretize Neural Ordinary Differential Equations?Michael E. Sander, Pierre Ablin, Gabriel PeyréNeurIPS 2022 · 42 citations
