Weight-Space Linear Recurrent Neural Networks
Roussel Desmond Nzoyem, Nawid Keshtmand, Enrique Crespo-Fernandez, Idriss Tsayem, Raúl Santos-Rodríguez, David Barton, Tom Deakin
Abstract
We introduce WARP (Weight-space Adaptive Recurrent Prediction), a simple yet powerful model that unifies weight-space learning with linear recurrence to redefine sequence modeling. Unlike conventional recurrent neural networks (RNNs) which collapse temporal dynamics into fixed-dimensional hidden states, WARP explicitly parametrizes its hidden state as the weights and biases of a distinct auxiliary neural network, and uses input differences to drive its recurrence. This brain-inspired formulation enables efficient gradient-free adaptation of the auxiliary network at test-time, in-context learning abilities, and seamless integration of domain-specific physical priors. Empirical validation shows that WARP matches or surpasses state-of-the-art baselines on diverse classification tasks, featuring in the top three in 4 out of 6 real-world challenging datasets. Furthermore, extensive experiments across sequential image completion, multivariate time series forecasting, and dynamical system reconstruction demonstrate its expressiveness and generalization capabilities. Remarkably, a physics-informed variant of our model outperforms the next best model by more than 10x. Ablation studies confirm the architectural necessity of key components, solidifying weight-space linear RNNs as a transformative paradigm for adaptive machine intelligence.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5ac9ac8b-036d-4195-89bf-03a057cd8a94Builds on45
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Informer: Beyond Efficient Transformer for Long Sequence Time-Series ForecastingHaoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang et al.AAAI 2021 · 7,289 citations
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- GMAN: A Graph Multi-Attention Network for Traffic PredictionChuanpan Zheng, Xiaoliang Fan, Cheng Wang, Jianzhong QiAAAI 2020 · 1,858 citations
Related papers
- Neural Wave Machines: Learning Spatiotemporally Structured Representations with Locally Coupled Oscillatory Recurrent Neural NetworksT. Anderson Keller, Max WellingICML 2023 · 26 citations
- Traveling Waves Encode The Recent Past and Enhance Sequence LearningT. Anderson Keller, Lyle Muller, Terrence J. Sejnowski, Max WellingICLR 2024 · 26 citations
- WAVE: Weighted Autoregressive Varying Gate for Time Series ForecastingJiecheng Lu, Xu Han, Yan Sun, Shihao YangICML 2025
- Learning long range dependencies through time reversal symmetry breakingGuillaume Pourcel, Maxence ErnoultNeurIPS 2025 · 9 citations
- Effectively Modeling Time Series with Simple Discrete State SpacesMichael Zhang, Khaled Kamal Saab, Michael Poli, Tri Dao et al.ICLR 2023 · 14 citations
