Weight-Space Linear Recurrent Neural Networks
Roussel Desmond Nzoyem, Nawid Keshtmand, Enrique Crespo-Fernandez, Idriss Tsayem, Raúl Santos-Rodríguez, David Barton, Tom Deakin
摘要
We introduce WARP (Weight-space Adaptive Recurrent Prediction), a simple yet powerful model that unifies weight-space learning with linear recurrence to redefine sequence modeling. Unlike conventional recurrent neural networks (RNNs) which collapse temporal dynamics into fixed-dimensional hidden states, WARP explicitly parametrizes its hidden state as the weights and biases of a distinct auxiliary neural network, and uses input differences to drive its recurrence. This brain-inspired formulation enables efficient gradient-free adaptation of the auxiliary network at test-time, in-context learning abilities, and seamless integration of domain-specific physical priors. Empirical validation shows that WARP matches or surpasses state-of-the-art baselines on diverse classification tasks, featuring in the top three in 4 out of 6 real-world challenging datasets. Furthermore, extensive experiments across sequential image completion, multivariate time series forecasting, and dynamical system reconstruction demonstrate its expressiveness and generalization capabilities. Remarkably, a physics-informed variant of our model outperforms the next best model by more than 10x. Ablation studies confirm the architectural necessity of key components, solidifying weight-space linear RNNs as a transformative paradigm for adaptive machine intelligence.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper45
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Informer: Beyond Efficient Transformer for Long Sequence Time-Series ForecastingHaoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang 等AAAI 2021 · 被引用 7,289 次
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 被引用 3,482 次
- GMAN: A Graph Multi-Attention Network for Traffic PredictionChuanpan Zheng, Xiaoliang Fan, Cheng Wang, Jianzhong QiAAAI 2020 · 被引用 1,858 次
相关 Paper
- Neural Wave Machines: Learning Spatiotemporally Structured Representations with Locally Coupled Oscillatory Recurrent Neural NetworksT. Anderson Keller, Max WellingICML 2023 · 被引用 26 次
- Traveling Waves Encode The Recent Past and Enhance Sequence LearningT. Anderson Keller, Lyle Muller, Terrence J. Sejnowski, Max WellingICLR 2024 · 被引用 26 次
- WAVE: Weighted Autoregressive Varying Gate for Time Series ForecastingJiecheng Lu, Xu Han, Yan Sun, Shihao YangICML 2025
- Learning long range dependencies through time reversal symmetry breakingGuillaume Pourcel, Maxence ErnoultNeurIPS 2025 · 被引用 9 次
- Effectively Modeling Time Series with Simple Discrete State SpacesMichael Zhang, Khaled Kamal Saab, Michael Poli, Tri Dao 等ICLR 2023 · 被引用 14 次
