Attention Approximates Sparse Distributed Memory
Trenton Bricken, Cengiz Pehlevan
摘要
While Attention has come to be an important mechanism in deep learning, there remains limited intuition for why it works so well. Here, we show that Transformer Attention can be closely related under certain data conditions to Kanerva's Sparse Distributed Memory (SDM), a biologically plausible associative memory model. We confirm that these conditions are satisfied in pre-trained GPT2 Transformer models. We discuss the implications of the Attention-SDM map and provide new computational and biological interpretations of Attention. 1 This cerebellar relationship is additionally compelling by the fact that cerebellum-like neuroanatomy exists in many other organisms including numerous insects (eg. the Drosophila Mushroom Body) and potentially cephalopods [17, 18, 19, 20, 21] . 35th Conference on Neural Information Processing Systems (NeurIPS 2021).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Birth of a Transformer: A Memory ViewpointAlberto Bietti, Vivien Cabannes, Diane Bouchacourt, Hervé Jégou 等NeurIPS 2023 · 被引用 182 次
- Universal Hopfield Networks: A General Framework for Single-Shot Associative Memory ModelsBeren Millidge, Tommaso Salvatori, Yuhang Song, Thomas Lukasiewicz 等ICML 2022 · 被引用 72 次
- Decoupled Context Processing for Context Augmented Language ModelingZonglin Li, Ruiqi Guo, Sanjiv KumarNeurIPS 2022 · 被引用 31 次
- Neuroformer: Multimodal and Multitask Generative Pretraining for Brain DataAntonis Antoniades, Yiyi Yu, Joseph Canzano, William Yang Wang 等ICLR 2024 · 被引用 21 次
- Do LLMs dream of elephants (when told not to)? Latent concept association and associative memory in transformersYibo Jiang, Goutham Rajendran, Pradeep Ravikumar, Bryon AragamNeurIPS 2024 · 被引用 19 次
它引用的顶会 Paper10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski 等USENIX Security 2021 · 被引用 2,866 次
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 被引用 2,665 次
- Generative Pretraining From PixelsMark Chen, Alec Radford, Rewon Child, Jeffrey Wu 等ICML 2020 · 被引用 1,773 次
相关 Paper
- Sparse Distributed Memory is a Continual LearnerTrenton Bricken, Xander Davies, Deepak Singh, Dmitry Krotov 等ICLR 2023 · 被引用 5 次
- Associative TransformerYuwei Sun, Hideya Ochiai, Zhirong Wu, Stephen Lin 等CVPR 2025
- Attention as Implicit Structural InferenceRyan Singh, Christopher L. BuckleyNeurIPS 2023 · 被引用 12 次
- On the Role of Hidden States of Modern Hopfield Network in TransformerTsubasa Masumura, Masato TakiNeurIPS 2025 · 被引用 2 次
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie 等NeurIPS 2020 · 被引用 3,159 次
