Second-order forward-mode optimization of recurrent neural networks for neuroscience
Youjing Yu, Rui Xia, Qingxi Ma, Máté Lengyel, Guillaume Hennequin
摘要
Training recurrent neural networks (RNNs) to perform neuroscience tasks can be challenging. Unlike in machine learning where any architectural modification of an RNN (e.g. GRU or LSTM) is acceptable if it facilitates training, the RNN models trained as models of brain dynamics are subject to plausibility constraints that fundamentally exclude the usual machine learning hacks. The “vanilla” RNNs commonly used in computational neuroscience find themselves plagued by ill-conditioned loss surfaces that complicate training and significantly hinder our capacity to investigate the brain dynamics underlying complex tasks. Moreover, some tasks may require very long time horizons which backpropagation cannot handle given typical GPU memory limits. Here, we develop SOFO, a second-order optimizer that efficiently navigates loss surfaces whilst not requiring backpropagation. By relying instead on easily parallelized batched forward-mode differentiation, SOFO enjoys constant memory cost in time. Moreover, unlike most second-order optimizers which involve inherently sequential operations, SOFO’s effective use of GPU parallelism yields a per-iteration wallclock time essentially on par with first-order gradient-based optimizers. We show vastly superior performance compared to Adam on a number of RNN tasks, including a difficult double-reaching motor task and the learning of an adaptive Kalman filter algorithm trained over a long horizon.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- RNNs perform task computations by dynamically warping neural representationsArthur Pellegrino, Angus ChadwickNeurIPS 2025 · 被引用 5 次
- Global curvature for second-order optimization of neural networksAlberto BernacchiaICML 2025
它引用的顶会 Paper11
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer 等NeurIPS 2021 · 被引用 3,862 次
- On the difficulty of learning chaotic dynamics with RNNsJonas M. Mikhaeil, Zahra Monfared, Daniel DurstewitzNeurIPS 2022 · 被引用 109 次
- Stochastic Subspace Cubic Newton MethodFilip Hanzely, Nikita Doikov, Yurii E. Nesterov, Peter RichtárikICML 2020 · 被引用 62 次
- Inferring Latent Dynamics Underlying Neural Population Activity via Neural Differential EquationsTimothy Doyeon Kim, Thomas Zhihao Luo, Jonathan W. Pillow, Carlos D. BrodyICML 2021 · 被引用 62 次
- Targeted Neural Dynamical ModelingCole L. Hurwitz, Akash Srivastava, Kai Xu, Justin Jude 等NeurIPS 2021 · 被引用 55 次
相关 Paper
- Training biologically plausible recurrent neural networks on cognitive tasks with long-term dependenciesWayne Soo, Vishwa Goudar, Xiao-Jing WangNeurIPS 2023 · 被引用 16 次
- Second-Order Neural ODE OptimizerGuan-Horng Liu, Tianrong Chen, Evangelos A. TheodorouNeurIPS 2021 · 被引用 20 次
- Understanding and Improving Optimization in Predictive Coding NetworksNicholas Alonso, Jeffrey L. Krichmar, Emre NeftciAAAI 2024 · 被引用 12 次
- Gradient Flossing: Improving Gradient Descent through Dynamic Control of JacobiansRainer EngelkenNeurIPS 2023 · 被引用 14 次
- Understanding and improving Shampoo and SOAP via Kullback-Leibler MinimizationWu Lin, Scott C. Lowe, Felix Dangel, Runa Eschenhagen 等ICLR 2026 · 被引用 15 次
