Neural Differential Equations for Learning to Program Neural Nets Through Continuous Learning Rules
Kazuki Irie, Francesco Faccio, Jürgen Schmidhuber
摘要
Neural ordinary differential equations (ODEs) have attracted much attention as continuous-time counterparts of deep residual neural networks (NNs), and numerous extensions for recurrent NNs have been proposed. Since the 1980s, ODEs have also been used to derive theoretical results for NN learning rules, e.g., the famous connection between Oja's rule and principal component analysis. Such rules are typically expressed as additive iterative update processes which have straightforward ODE counterparts. Here we introduce a novel combination of learning rules and Neural ODEs to build continuous-time sequence processing nets that learn to manipulate short-term memory in rapidly changing synaptic connections of other nets. This yields continuous-time counterparts of Fast Weight Programmers and linear Transformers. Our novel models outperform the best existing Neural Controlled Differential Equation based models on various time series classification tasks, while also addressing their fundamental scalability limitations. Our code is public. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Nested Learning: The Illusion of Deep Learning ArchitecturesAli Behrouz, Meisam Razaviyayn, Peilin Zhong, Vahab MirrokniNeurIPS 2025 · 被引用 96 次
- DualDynamics: Synergizing Implicit and Explicit Methods for Robust Irregular Time Series AnalysisYongKyung Oh, Dong-Young Lim, Sungil KimAAAI 2025 · 被引用 5 次
- Controlled Differential Equations on Long Sequences via Non-standard WaveletsSourav Pal, Zhanpeng Zeng, Sathya N. Ravi, Vikas SinghICML 2023 · 被引用 3 次
- Images as Weight Matrices: Sequential Image Generation Through Synaptic Learning RulesKazuki Irie, Jürgen SchmidhuberICLR 2023 · 被引用 1 次
- Learning to (Learn at Test Time): RNNs with Expressive Hidden StatesYu Sun, Xinhao Li, Karan Dalal, Jiarui Xu 等ICML 2025
它引用的顶会 Paper13
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 被引用 2,665 次
- Neural Controlled Differential Equations for Irregular Time SeriesPatrick Kidger, James Morrill, James Foster, Terry J. LyonsNeurIPS 2020 · 被引用 850 次
- Linear Transformers Are Secretly Fast Weight ProgrammersImanol Schlag, Kazuki Irie, Jürgen SchmidhuberICML 2021 · 被引用 394 次
- Neural Rough Differential Equations for Long Time SeriesJames Morrill, Cristopher Salvi, Patrick Kidger, James FosterICML 2021 · 被引用 176 次
- Going Beyond Linear Transformers with Recurrent Fast Weight ProgrammersKazuki Irie, Imanol Schlag, Róbert Csordás, Jürgen SchmidhuberNeurIPS 2021 · 被引用 101 次
相关 Paper
- How Deep Do We Need: Accelerating Training and Inference of Neural ODEs via Control PerspectiveKeyan Miao, Konstantinos GatsisICML 2024 · 被引用 2 次
- ControlSynth Neural ODEs: Modeling Dynamical Systems with Guaranteed ConvergenceWenjie Mei, Dongzhe Zheng, Shihua LiNeurIPS 2024 · 被引用 20 次
- Generalization bounds for neural ordinary differential equations and deep residual networksPierre MarionNeurIPS 2023 · 被引用 37 次
- Neural Dynamics on Complex NetworksChengxi Zang, Fei WangKDD 2020 · 被引用 4 次
- Learning Efficient and Robust Ordinary Differential Equations via Invertible Neural NetworksWeiming Zhi, Tin Lai, Lionel Ott, Edwin V. Bonilla 等ICML 2022 · 被引用 26 次
