Adaptive Checkpoint Adjoint Method for Gradient Estimation in Neural ODE
Juntang Zhuang, Nicha C. Dvornek, Xiaoxiao Li, Sekhar Tatikonda, Xenophon Papademetris, James S. Duncan
摘要
Neural ordinary differential equations (NODEs) have recently attracted increasing attention; however, their empirical performance on benchmark tasks (e.g. image classification) are significantly inferior to discrete-layer models. We demonstrate an explanation for their poorer performance is the inaccuracy of existing gradient estimation methods: the adjoint method has numerical errors in reverse-mode integration; the naive method directly back-propagates through ODE solvers, but suffers from a redundantly deep computation graph when searching for the optimal stepsize. We propose the Adaptive Checkpoint Adjoint (ACA) method: in automatic differentiation, ACA applies a trajectory checkpoint strategy which records the forward-mode trajectory as the reverse-mode trajectory to guarantee accuracy; ACA deletes redundant components for shallow computation graphs; and ACA supports adaptive solvers. On image classification tasks, compared with the adjoint and naive method, ACA achieves half the error rate in half the training time; NODE trained with ACA outperforms ResNet in both accuracy and test-retest reliability. On time-series modeling, ACA outperforms competing methods. Finally, in an example of the three-body problem, we show NODE with ACA can incorporate physical knowledge to achieve better accuracy. We provide the PyTorch implementation of ACA: https://github.com/ juntang-zhuang/torch-ACA .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper41
- Liquid Time-constant NetworksRamin M. Hasani, Mathias Lechner, Alexander Amini, Daniela Rus 等AAAI 2021 · 被引用 399 次
- Neural Flows: Efficient Alternative to Neural ODEsMarin Bilos, Johanna Sommer, Syama Sundar Rangapuram, Tim Januschowski 等NeurIPS 2021 · 被引用 151 次
- SE(3) Equivariant Graph Neural Networks with Complete Local FramesWeitao Du, He Zhang, Yuanqi Du, Qi Meng 等ICML 2022 · 被引用 111 次
- Robust Evaluation of Diffusion-Based Adversarial PurificationMinjong Lee, Dongwoo KimICCV 2023 · 被引用 96 次
- Machine learning structure preserving brackets for forecasting irreversible processesKookjin Lee, Nathaniel Trask, Panos StinisNeurIPS 2021 · 被引用 80 次
相关 Paper
- MALI: A memory efficient and reverse accurate integrator for Neural ODEsJuntang Zhuang, Nicha C. Dvornek, Sekhar Tatikonda, James S. DuncanICLR 2021 · 被引用 61 次
- Improving Neural Ordinary Differential Equations with Nesterov's Accelerated Gradient MethodHo Huu Nghia Nguyen, Tan Nguyen, Huyen Vo, Stanley J. Osher 等NeurIPS 2022 · 被引用 28 次
- Symplectic Adjoint Method for Exact Gradient of Neural ODE with Minimal MemoryTakashi Matsubara, Yuto Miyatake, Takaharu YaguchiNeurIPS 2021 · 被引用 30 次
- Heavy Ball Neural Ordinary Differential EquationsHedi Xia, Vai Suliafu, Hangjie Ji, Tan M. Nguyen 等NeurIPS 2021 · 被引用 75 次
- Do Residual Neural Networks discretize Neural Ordinary Differential Equations?Michael E. Sander, Pierre Ablin, Gabriel PeyréNeurIPS 2022 · 被引用 42 次
