Transformer as a hippocampal memory consolidation model based on NMDAR-inspired nonlinearity
Dong Kyum Kim, Jea Kwon, Meeyoung Cha, Chul Lee
Abstract
The hippocampus plays a critical role in learning, memory, and spatial representation, processes that depend on the NMDA receptor (NMDAR). Inspired by recent findings that compare deep learning models to the hippocampus, we propose a new nonlinear activation function that mimics NMDAR dynamics. NMDAR-like nonlinearity shifts short-term working memory into long-term reference memory in transformers, thus enhancing a process that is similar to memory consolidation in the mammalian brain. We design a navigation task assessing these two memory functions and show that manipulating the activation function (i.e., mimicking the Mg 2+ -gating of NMDAR) disrupts long-term memory processes. Our experiments suggest that place cell-like functions and reference memory reside in the feed-forward network layer of transformers and that nonlinearity drives these processes. We discuss the role of NMDAR-like nonlinearity in establishing this striking resemblance between transformer architecture and hippocampal spatial representation. * Equal contribution. † Corresponding authors. 37th Conference on Neural Information Processing Systems (NeurIPS 2023). Methods Designing a 2D navigation task to test the role of working memory and reference memory We designed a sensory observation prediction task in which an agent randomly walks in a 2D grid environment and is trained to predict subsequent sensory observations (see Fig. 2a ) [12] . The agent
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- NeurIPS: Neuro-anatomical Inductive Priors for Sphere-based Brain DecodingSijin Yu, Zijiao Chen, Zhenyu Yang, Zihao Tan et al.ICML 2026
- SIM: Surface-based fMRI Analysis for Inter-Subject Multimodal Decoding from Movie-Watching ExperimentsSimon Dahan, Gabriel Bénédict, Logan Zane John Williams, Yourong Guo et al.ICLR 2025
Builds on12
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
- Hopfield Networks is All You NeedHubert Ramsauer, Bernhard Schäfl, Johannes Lehner, Philipp Seidl et al.ICLR 2021 · 620 citations
- Linear Transformers Are Secretly Fast Weight ProgrammersImanol Schlag, Kazuki Irie, Jürgen SchmidhuberICML 2021 · 394 citations
Related papers
- Relating transformers to models and neural representations of the hippocampal formationJames C. R. Whittington, Joseph Warren, Tim E. J. BehrensICLR 2022 · 110 citations
- A Cognitive Model for Learning Abstract Relational Structures from Memory-based Decision-Making TasksHaruo HosoyaICLR 2024 · 1 citation
- Hippoformer: Integrating Hippocampus-inspired Spatial Memory with TransformersTiantian Li, Xingxing Cao, Yifei Wang, Xiaojiao Yang et al.ICLR 2026
- Emergence of Spatial Representation in an Actor-Critic Agent with Hippocampus-Inspired Sequence GeneratorXiao-Xiong Lin, Yuk Hoi Yiu, Christian LeiboldICLR 2026 · 2 citations
- Time Makes Space: Emergence of Place Fields in Networks Encoding Temporally Continuous Sensory ExperiencesZhaoze Wang, Ronald W. Di Tullio, Spencer Rooke, Vijay BalasubramanianNeurIPS 2024 · 16 citations
