Stateful ODE-Nets using Basis Function Expansions
Alejandro F. Queiruga, N. Benjamin Erichson, Liam Hodgkinson, Michael W. Mahoney
摘要
The recently-introduced class of ordinary differential equation networks (ODE-Nets) establishes a fruitful connection between deep learning and dynamical systems. In this work, we reconsider formulations of the weights as continuous-in-depth functions using linear combinations of basis functions which enables us to leverage parameter transformations such as function projections. In turn, this view allows us to formulate a novel stateful ODE-Block that handles stateful layers. The benefits of this new ODE-Block are twofold: first, it enables incorporating meaningful continuous-in-depth batch normalization layers to achieve state-of-the-art performance; second, it enables compressing the weights through a change of basis, without retraining, while maintaining near state-of-the-art performance and reducing both inference time and memory footprint. Performance is demonstrated by applying our stateful ODE-Block to (a) image classification tasks using convolutional units and (b) sentence-tagging tasks using transformer encoder units.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Learning Physics Constrained Dynamics Using AutoencodersTsung-Yen Yang, Justinian Rosca, Karthik Narasimhan, Peter J. RamadgeNeurIPS 2022 · 被引用 39 次
- Generalization bounds for neural ordinary differential equations and deep residual networksPierre MarionNeurIPS 2023 · 被引用 37 次
- Implicit regularization of deep residual networks towards neural ODEsPierre Marion, Yu-Han Wu, Michael Eli Sander, Gérard BiauICLR 2024 · 被引用 24 次
它引用的顶会 Paper11
- Reducing Transformer Depth on Demand with Structured DropoutAngela Fan, Edouard Grave, Armand JoulinICLR 2020 · 被引用 695 次
- Symplectic ODE-Net: Learning Hamiltonian Dynamics with ControlYaofeng Desmond Zhong, Biswadip Dey, Amit ChakrabortyICLR 2020 · 被引用 319 次
- Depth-Adaptive TransformerMaha Elbayad, Jiatao Gu, Edouard Grave, Michael AuliICLR 2020 · 被引用 264 次
- Dissecting Neural ODEsStefano Massaroli, Michael Poli, Jinkyoo Park, Atsushi Yamashita 等NeurIPS 2020 · 被引用 261 次
- Adaptive Checkpoint Adjoint Method for Gradient Estimation in Neural ODEJuntang Zhuang, Nicha C. Dvornek, Xiaoxiao Li, Sekhar Tatikonda 等ICML 2020 · 被引用 125 次
相关 Paper
- Improving Neural ODE Training with Temporal Adaptive Batch NormalizationSu Zheng, Zhengqi Gao, Fan-Keng Sun, Duane S. Boning 等NeurIPS 2024 · 被引用 5 次
- Sparse Flows: Pruning Continuous-depth ModelsLucas Liebenwein, Ramin M. Hasani, Alexander Amini, Daniela RusNeurIPS 2021 · 被引用 21 次
- Neural Differential Equations for Learning to Program Neural Nets Through Continuous Learning RulesKazuki Irie, Francesco Faccio, Jürgen SchmidhuberNeurIPS 2022 · 被引用 24 次
- A shooting formulation of deep learningFrançois-Xavier Vialard, Roland Kwitt, Susan Wei, Marc NiethammerNeurIPS 2020 · 被引用 16 次
- CNM-UNet: Continuous Ordinary Differential Equations for Medical Image SegmentationTianqi Xu, Yashi Zhu, Quansong He, Yue Cao 等AAAI 2026
