Linear Transformers Are Secretly Fast Weight Programmers
Imanol Schlag, Kazuki Irie, Jürgen Schmidhuber
摘要
We show the formal equivalence of linearised self-attention mechanisms and fast weight controllers from the early '90s, where a slow"neural net learns by gradient descent to program the fast weights"of another net through sequences of elementary programming instructions which are additive outer products of self-invented activation patterns (today called keys and values). Such Fast Weight Programmers (FWPs) learn to manipulate the contents of a finite memory and dynamically interact with it. We infer a memory capacity limitation of recent linearised softmax attention variants, and replace the purely additive outer products by a delta rule-like programming instruction, such that the FWP can more easily learn to correct the current mapping from keys to values. The FWP also learns to compute dynamically changing learning rates. We also propose a new kernel function to linearise attention which balances simplicity and effectiveness. We conduct experiments on synthetic retrieval problems as well as standard machine translation and language modelling tasks which demonstrate the benefits of our methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper177
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space DualityTri Dao, Albert GuICML 2024 · 被引用 1,407 次
- Transformers Learn In-Context by Gradient DescentJohannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento 等ICML 2023 · 被引用 729 次
- xLSTM: Extended Long Short-Term MemoryMaximilian Beck, Korbinian Pöppel, Markus Spanring, Andreas Auer 等NeurIPS 2024 · 被引用 703 次
- Choose a Transformer: Fourier or GalerkinShuhao CaoNeurIPS 2021 · 被引用 516 次
- Hyena Hierarchy: Towards Larger Convolutional Language ModelsMichael Poli, Stefano Massaroli, Eric Nguyen, Daniel Y. Fu 等ICML 2023 · 被引用 481 次
它引用的顶会 Paper9
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 被引用 2,665 次
- Long Range Arena : A Benchmark for Efficient TransformersYi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen 等ICLR 2021 · 被引用 881 次
- Compressive Transformers for Long-Range Sequence ModellingJack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, Chloe Hillier 等ICLR 2020 · 被引用 833 次
- Hopfield Networks is All You NeedHubert Ramsauer, Bernhard Schäfl, Johannes Lehner, Philipp Seidl 等ICLR 2021 · 被引用 620 次
- Random Feature AttentionHao Peng, Nikolaos Pappas, Dani Yogatama, Roy Schwartz 等ICLR 2021 · 被引用 425 次
相关 Paper
- Going Beyond Linear Transformers with Recurrent Fast Weight ProgrammersKazuki Irie, Imanol Schlag, Róbert Csordás, Jürgen SchmidhuberNeurIPS 2021 · 被引用 101 次
- Blending Complementary Memory Systems in Hybrid Quadratic-Linear TransformersKazuki Irie, Morris Yau, Samuel J. GershmanNeurIPS 2025 · 被引用 13 次
- A Modern Self-Referential Weight Matrix That Learns to Modify ItselfKazuki Irie, Imanol Schlag, Róbert Csordás, Jürgen SchmidhuberICML 2022 · 被引用 42 次
- Images as Weight Matrices: Sequential Image Generation Through Synaptic Learning RulesKazuki Irie, Jürgen SchmidhuberICLR 2023 · 被引用 1 次
- Rethinking Attention with PerformersKrzysztof Marcin Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song 等ICLR 2021 · 被引用 122 次
