Learning Useful Representations of Recurrent Neural Network Weight Matrices
Vincent Herrmann, Francesco Faccio, Jürgen Schmidhuber
摘要
Recurrent Neural Networks (RNNs) are general-purpose parallel-sequential computers. The program of an RNN is its weight matrix. How to learn useful representations of RNN weights that facilitate RNN analysis as well as downstream tasks? While the mechanistic approach directly looks at some RNN's weights to predict its behavior, the functionalist approach analyzes its overall functionality-specifically, its input-output mapping. We consider several mechanistic approaches for RNN weights and adapt the permutation equivariant Deep Weight Space layer for RNNs. Our two novel functionalist approaches extract information from RNN weights by 'interrogating' the RNN through probing inputs. We develop a theoretical framework that demonstrates conditions under which the functionalist approach can generate rich representations that help determine RNN behavior. We release the first two 'model zoo' datasets for RNN weight representation learning. One consists of generative models of a class of formal languages, and the other one of classifiers of sequentially processed MNIST digits. With the help of an emulation-based self-supervised learning technique we compare and evaluate the different RNN weight encoding techniques on multiple downstream applications. On the most challenging one, namely predicting which exact task the RNN was trained on, functionalist approaches show clear superiority.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Improved Generalization of Weight Space Networks via AugmentationsAviv Shamsian, Aviv Navon, David W. Zhang, Yan Zhang 等ICML 2024 · 被引用 19 次
- GradMetaNet: An Equivariant Architecture for Learning on GradientsYoav Gelberg, Yam Eitan, Aviv Navon, Aviv Shamsian 等NeurIPS 2025 · 被引用 8 次
- What Linear Probes Miss: Multi-View Probing for Weight-Space LearningEunwoo Heo, Kyeongkook Seo, Jaejun YooICML 2026
- Deep Linear Probe Generators for Weight Space LearningJonathan Kahana, Eliahu Horwitz, Imri Shuval, Yedid HoshenICLR 2025
- Parameter Manifold PurificationJiacong Hu, Jinxun Wu, Shengxuming Zhang, Shunyu Liu 等ICML 2026
它引用的顶会 Paper14
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 被引用 3,482 次
- Are Transformers universal approximators of sequence-to-sequence functions?Chulhee Yun, Srinadh Bhojanapalli, Ankit Singh Rawat, Sashank J. Reddi 等ICLR 2020 · 被引用 481 次
- From data to functa: Your data point is a function and you can treat it like oneEmilien Dupont, Hyunjik Kim, S. M. Ali Eslami, Danilo Jimenez Rezende 等ICML 2022 · 被引用 209 次
- Socratic Models: Composing Zero-Shot Multimodal Reasoning with LanguageAndy Zeng, Maria Attarian, Brian Ichter, Krzysztof Marcin Choromanski 等ICLR 2023 · 被引用 171 次
- Equivariant Architectures for Learning in Deep Weight SpacesAviv Navon, Aviv Shamsian, Idan Achituve, Ethan Fetaya 等ICML 2023 · 被引用 101 次
相关 Paper
- Permutation Equivariant Neural FunctionalsAllan Zhou, Kaien Yang, Kaylee Burns, Adriano Cardace 等NeurIPS 2023 · 被引用 84 次
- Decision-Guided Weighted Automata Extraction from Recurrent Neural NetworksXiyue Zhang, Xiaoning Du, Xiaofei Xie, Lei Ma 等AAAI 2021 · 被引用 25 次
- Universal Neural FunctionalsAllan Zhou, Chelsea Finn, James HarrisonNeurIPS 2024 · 被引用 27 次
- A unified theory of feature learning in RNNs and DNNsJan Bauer, Kirsten Fischer, Moritz Helias, Agostina PalmigianoICML 2026 · 被引用 4 次
- Neural Functional TransformersAllan Zhou, Kaien Yang, Yiding Jiang, Kaylee Burns 等NeurIPS 2023 · 被引用 53 次
