Neural Differential Equations for Learning to Program Neural Nets Through Continuous Learning Rules
Kazuki Irie, Francesco Faccio, Jürgen Schmidhuber
Abstract
Neural ordinary differential equations (ODEs) have attracted much attention as continuous-time counterparts of deep residual neural networks (NNs), and numerous extensions for recurrent NNs have been proposed. Since the 1980s, ODEs have also been used to derive theoretical results for NN learning rules, e.g., the famous connection between Oja's rule and principal component analysis. Such rules are typically expressed as additive iterative update processes which have straightforward ODE counterparts. Here we introduce a novel combination of learning rules and Neural ODEs to build continuous-time sequence processing nets that learn to manipulate short-term memory in rapidly changing synaptic connections of other nets. This yields continuous-time counterparts of Fast Weight Programmers and linear Transformers. Our novel models outperform the best existing Neural Controlled Differential Equation based models on various time series classification tasks, while also addressing their fundamental scalability limitations. Our code is public. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 19b1a66e-5c5e-473e-837c-46e8c64e2898Cited by top-tier papers6
- Nested Learning: The Illusion of Deep Learning ArchitecturesAli Behrouz, Meisam Razaviyayn, Peilin Zhong, Vahab MirrokniNeurIPS 2025 · 96 citations
- DualDynamics: Synergizing Implicit and Explicit Methods for Robust Irregular Time Series AnalysisYongKyung Oh, Dong-Young Lim, Sungil KimAAAI 2025 · 5 citations
- Controlled Differential Equations on Long Sequences via Non-standard WaveletsSourav Pal, Zhanpeng Zeng, Sathya N. Ravi, Vikas SinghICML 2023 · 3 citations
- Images as Weight Matrices: Sequential Image Generation Through Synaptic Learning RulesKazuki Irie, Jürgen SchmidhuberICLR 2023 · 1 citation
- Learning to (Learn at Test Time): RNNs with Expressive Hidden StatesYu Sun, Xinhao Li, Karan Dalal, Jiarui Xu et al.ICML 2025
Builds on13
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
- Neural Controlled Differential Equations for Irregular Time SeriesPatrick Kidger, James Morrill, James Foster, Terry J. LyonsNeurIPS 2020 · 850 citations
- Linear Transformers Are Secretly Fast Weight ProgrammersImanol Schlag, Kazuki Irie, Jürgen SchmidhuberICML 2021 · 394 citations
- Neural Rough Differential Equations for Long Time SeriesJames Morrill, Cristopher Salvi, Patrick Kidger, James FosterICML 2021 · 176 citations
- Going Beyond Linear Transformers with Recurrent Fast Weight ProgrammersKazuki Irie, Imanol Schlag, Róbert Csordás, Jürgen SchmidhuberNeurIPS 2021 · 101 citations
Related papers
- How Deep Do We Need: Accelerating Training and Inference of Neural ODEs via Control PerspectiveKeyan Miao, Konstantinos GatsisICML 2024 · 2 citations
- ControlSynth Neural ODEs: Modeling Dynamical Systems with Guaranteed ConvergenceWenjie Mei, Dongzhe Zheng, Shihua LiNeurIPS 2024 · 20 citations
- Generalization bounds for neural ordinary differential equations and deep residual networksPierre MarionNeurIPS 2023 · 37 citations
- Neural Dynamics on Complex NetworksChengxi Zang, Fei WangKDD 2020 · 4 citations
- Learning Efficient and Robust Ordinary Differential Equations via Invertible Neural NetworksWeiming Zhi, Tin Lai, Lionel Ott, Edwin V. Bonilla et al.ICML 2022 · 26 citations
