Stateful ODE-Nets using Basis Function Expansions
Alejandro F. Queiruga, N. Benjamin Erichson, Liam Hodgkinson, Michael W. Mahoney
Abstract
The recently-introduced class of ordinary differential equation networks (ODE-Nets) establishes a fruitful connection between deep learning and dynamical systems. In this work, we reconsider formulations of the weights as continuous-in-depth functions using linear combinations of basis functions which enables us to leverage parameter transformations such as function projections. In turn, this view allows us to formulate a novel stateful ODE-Block that handles stateful layers. The benefits of this new ODE-Block are twofold: first, it enables incorporating meaningful continuous-in-depth batch normalization layers to achieve state-of-the-art performance; second, it enables compressing the weights through a change of basis, without retraining, while maintaining near state-of-the-art performance and reducing both inference time and memory footprint. Performance is demonstrated by applying our stateful ODE-Block to (a) image classification tasks using convolutional units and (b) sentence-tagging tasks using transformer encoder units.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3e8777a0-9212-4d5e-9676-447a9483e1eaCited by top-tier papers3
- Learning Physics Constrained Dynamics Using AutoencodersTsung-Yen Yang, Justinian Rosca, Karthik Narasimhan, Peter J. RamadgeNeurIPS 2022 · 39 citations
- Generalization bounds for neural ordinary differential equations and deep residual networksPierre MarionNeurIPS 2023 · 37 citations
- Implicit regularization of deep residual networks towards neural ODEsPierre Marion, Yu-Han Wu, Michael Eli Sander, Gérard BiauICLR 2024 · 24 citations
Builds on11
- Reducing Transformer Depth on Demand with Structured DropoutAngela Fan, Edouard Grave, Armand JoulinICLR 2020 · 695 citations
- Symplectic ODE-Net: Learning Hamiltonian Dynamics with ControlYaofeng Desmond Zhong, Biswadip Dey, Amit ChakrabortyICLR 2020 · 319 citations
- Depth-Adaptive TransformerMaha Elbayad, Jiatao Gu, Edouard Grave, Michael AuliICLR 2020 · 264 citations
- Dissecting Neural ODEsStefano Massaroli, Michael Poli, Jinkyoo Park, Atsushi Yamashita et al.NeurIPS 2020 · 261 citations
- Adaptive Checkpoint Adjoint Method for Gradient Estimation in Neural ODEJuntang Zhuang, Nicha C. Dvornek, Xiaoxiao Li, Sekhar Tatikonda et al.ICML 2020 · 125 citations
Related papers
- Improving Neural ODE Training with Temporal Adaptive Batch NormalizationSu Zheng, Zhengqi Gao, Fan-Keng Sun, Duane S. Boning et al.NeurIPS 2024 · 5 citations
- Sparse Flows: Pruning Continuous-depth ModelsLucas Liebenwein, Ramin M. Hasani, Alexander Amini, Daniela RusNeurIPS 2021 · 21 citations
- Neural Differential Equations for Learning to Program Neural Nets Through Continuous Learning RulesKazuki Irie, Francesco Faccio, Jürgen SchmidhuberNeurIPS 2022 · 24 citations
- A shooting formulation of deep learningFrançois-Xavier Vialard, Roland Kwitt, Susan Wei, Marc NiethammerNeurIPS 2020 · 16 citations
- CNM-UNet: Continuous Ordinary Differential Equations for Medical Image SegmentationTianqi Xu, Yashi Zhu, Quansong He, Yue Cao et al.AAAI 2026
