Integrated Episodic and Semantic Memory via Modulating Transformer FeedForward Layers
Yiqun Yao, Xiang Li, Xin Jiang, Xuezhi Fang, Naitong Yu, Siwei Dong, Wenjia Ma, Jing Li, Aixin Sun, Yequan Wang
Abstract
It is widely recognized that, after generative pre-training, Transformer FeedForward layers implicitly function as semantic memory, encoding linguistic and factual knowledge, while the contexts in key–value (KV) cache contain raw events, serving as the source of models' episodic memory. In this work, we show that a same group of Transformer FeedForward-layer parameters can both be semantic and episodic memory, which is retrievable without explicitly attending to the related KV cache. To realize this idea, we introduce Hypermem, a hypernetwork that recurrently maps contexts into targeted updates of FeedForward parameters. We post-train the hypernetwork using continuation and random-access associative memory objectives, eliminating the need for test-time training. Extensive experiments demonstrate that our approach outperforms related methods, including MemoryLLM and generative adapter, on memory retrieval, long-context question answering, and personalization benchmarks, establishing a new state of the art for hypernetwork-based memory mechanisms. Our results suggest that directly bridging data and parameters provides a viable direction for exploring next-generation foundation models with more flexible and persistent memory capabilities.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e104ffcb-a89d-44e4-8c63-96b96a4db7d5Builds on18
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
- Titans: Learning to Memorize at Test TimeAli Behrouz, Peilin Zhong, Vahab MirrokniNeurIPS 2025 · 368 citations
Related papers
- EpMAN: Episodic Memory AttentioN for Generalizing to Longer ContextsSubhajit Chaudhury, Payel Das, Sarathkrishna Swaminathan, Georgios Kollias et al.ACL 2025
- HyperPrompt: Prompt-based Task-Conditioning of TransformersYun He, Huaixiu Steven Zheng, Yi Tay, Jai Prakash Gupta et al.ICML 2022 · 110 citations
- Pretraining with hierarchical memories: separating long-tail and common knowledgeHadi Pouransari, David Grangier, C Thomas, Michael Kirchhof et al.ICLR 2026 · 11 citations
- GradMem: Learning to Write Context into Memory with Test-Time Gradient DescentYuri Kuratov, Matvey Kairov, Aydar Bulatov, Ivan Rodkin et al.ICML 2026 · 3 citations
- HyperMem: Hypergraph Memory for Long-Term ConversationsJuwei Yue, Chuanrui Hu, Jiawei Sheng, Zuyi Zhou et al.ACL 2026 · 4 citations
