Attention Approximates Sparse Distributed Memory
Trenton Bricken, Cengiz Pehlevan
Abstract
While Attention has come to be an important mechanism in deep learning, there remains limited intuition for why it works so well. Here, we show that Transformer Attention can be closely related under certain data conditions to Kanerva's Sparse Distributed Memory (SDM), a biologically plausible associative memory model. We confirm that these conditions are satisfied in pre-trained GPT2 Transformer models. We discuss the implications of the Attention-SDM map and provide new computational and biological interpretations of Attention. 1 This cerebellar relationship is additionally compelling by the fact that cerebellum-like neuroanatomy exists in many other organisms including numerous insects (eg. the Drosophila Mushroom Body) and potentially cephalopods [17, 18, 19, 20, 21] . 35th Conference on Neural Information Processing Systems (NeurIPS 2021).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 75ce570f-a288-4eec-8465-25fdbd74b6e0Cited by top-tier papers14
- Birth of a Transformer: A Memory ViewpointAlberto Bietti, Vivien Cabannes, Diane Bouchacourt, Hervé Jégou et al.NeurIPS 2023 · 182 citations
- Universal Hopfield Networks: A General Framework for Single-Shot Associative Memory ModelsBeren Millidge, Tommaso Salvatori, Yuhang Song, Thomas Lukasiewicz et al.ICML 2022 · 72 citations
- Decoupled Context Processing for Context Augmented Language ModelingZonglin Li, Ruiqi Guo, Sanjiv KumarNeurIPS 2022 · 31 citations
- Neuroformer: Multimodal and Multitask Generative Pretraining for Brain DataAntonis Antoniades, Yiyi Yu, Joseph Canzano, William Yang Wang et al.ICLR 2024 · 21 citations
- Do LLMs dream of elephants (when told not to)? Latent concept association and associative memory in transformersYibo Jiang, Goutham Rajendran, Pradeep Ravikumar, Bryon AragamNeurIPS 2024 · 19 citations
Builds on10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski et al.USENIX Security 2021 · 2,866 citations
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
- Generative Pretraining From PixelsMark Chen, Alec Radford, Rewon Child, Jeffrey Wu et al.ICML 2020 · 1,773 citations
Related papers
- Sparse Distributed Memory is a Continual LearnerTrenton Bricken, Xander Davies, Deepak Singh, Dmitry Krotov et al.ICLR 2023 · 5 citations
- Associative TransformerYuwei Sun, Hideya Ochiai, Zhirong Wu, Stephen Lin et al.CVPR 2025
- Attention as Implicit Structural InferenceRyan Singh, Christopher L. BuckleyNeurIPS 2023 · 12 citations
- On the Role of Hidden States of Modern Hopfield Network in TransformerTsubasa Masumura, Masato TakiNeurIPS 2025 · 2 citations
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
