A shooting formulation of deep learning
François-Xavier Vialard, Roland Kwitt, Susan Wei, Marc Niethammer
Abstract
Continuous-depth neural networks can be viewed as deep limits of discrete neural networks whose dynamics resemble a discretization of an ordinary differential equation (ODE). Although important steps have been taken to realize the advantages of such continuous formulations, most current techniques are not truly continuous-depth as they assume identical layers. Indeed, existing works throw into relief the myriad difficulties presented by an infinite-dimensional parameter space in learning a continuous-depth neural ODE. To this end, we introduce a shooting formulation which shifts the perspective from parameterizing a network layer-by-layer to parameterizing over optimal networks described only by a set of initial conditions. For scalability, we propose a novel particle-ensemble parametrization which fully specifies the optimal weight trajectory of the continuous-depth neural network. Our experiments show that our particle-ensemble shooting formulation can achieve competitive performance, especially on long-range forecasting tasks. Finally, though the current work is inspired by continuous-depth neural networks, the particle-ensemble shooting formulation also applies to discrete-time networks and may lead to a new fertile area of research in deep learning parametrization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 31df249c-c0aa-49f0-ba0d-ca4b92cb3810Cited by top-tier papers2
- On global convergence of ResNets: From finite to infinite width using linear parameterizationRaphaël Barboni, Gabriel Peyré, François-Xavier VialardNeurIPS 2022 · 14 citations
- Imbedding Deep Neural NetworksAndrew Corbett, Dmitry KanginICLR 2022 · 2 citations
Builds on3
- Hamiltonian Generative NetworksPeter Toth, Danilo J. Rezende, Andrew Jaegle, Sébastien Racanière et al.ICLR 2020 · 242 citations
- Approximation Capabilities of Neural ODEs and Invertible Residual NetworksHan Zhang, Xi Gao, Jacob Unterman, Tom ArodzICML 2020 · 114 citations
- How to Train Your Neural ODE: the World of Jacobian and Kinetic RegularizationChris Finlay, Jörn-Henrik Jacobsen, Levon Nurbekyan, Adam M. ObermanICML 2020 · 76 citations
Related papers
- Sparse Flows: Pruning Continuous-depth ModelsLucas Liebenwein, Ramin M. Hasani, Alexander Amini, Daniela RusNeurIPS 2021 · 21 citations
- Dissecting Neural ODEsStefano Massaroli, Michael Poli, Jinkyoo Park, Atsushi Yamashita et al.NeurIPS 2020 · 261 citations
- Neural Piecewise-Constant Delay Differential EquationsQunxi Zhu, Yifei Shen, Dongsheng Li, Wei LinAAAI 2022 · 9 citations
- Anamnesic Neural Differential Equations with Orthogonal Polynomial ProjectionsEdward De Brouwer, Rahul G. KrishnanICLR 2023 · 2 citations
- Learning Efficient and Robust Ordinary Differential Equations via Invertible Neural NetworksWeiming Zhi, Tin Lai, Lionel Ott, Edwin V. Bonilla et al.ICML 2022 · 26 citations
