Understanding Factual Recall in Transformers via Associative Memories
Eshaan Nichani, Jason D. Lee, Alberto Bietti
摘要
Large language models have demonstrated an impressive ability to perform factual recall. Prior work has found that transformers trained on factual recall tasks can store information at a rate proportional to their parameter count. In our work, we show that shallow transformers can use a combination of associative memories to obtain such near optimal storage capacity. We begin by proving that the storage capacities of both linear and MLP associative memories scale linearly with parameter count. We next introduce a synthetic factual recall task, and prove that a transformer with a single layer of self-attention followed by an MLP can obtain 100% accuracy on the task whenever either the total number of self-attention parameters or MLP parameters scales (up to log factors) linearly with the number of facts. In particular, the transformer can trade off between using the value matrices or the MLP as an associative memory to store the dataset of facts. We complement these expressivity results with an analysis of the gradient flow trajectory of a simplified linear attention model trained on our factual recall task, where we show that the model exhibits sequential learning behavior.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper23
- Muon Outperforms Adam in Tail-End Associative Memory LearningShuche Wang, Fengzhuo Zhang, Jiaxiang Li, Cunxiao Du 等ICLR 2026 · 被引用 40 次
- Emergence and scaling laws in SGD learning of shallow neural networksYunwei Ren, Eshaan Nichani, Denny Wu, Jason D. LeeNeurIPS 2025 · 被引用 33 次
- Emergence of Superposition: Unveiling the Training Dynamics of Chain of Continuous ThoughtHanlin Zhu, Shibo Hao, Zhiting Hu, Jiantao Jiao 等ICLR 2026 · 被引用 21 次
- Why and How LLMs Hallucinate: Connecting the Dots with Subsequence AssociationsYiyou Sun, Yu Gai, Lijie Chen, Abhilasha Ravichander 等NeurIPS 2025 · 被引用 20 次
- Data Mixing Can Induce Phase Transitions in Knowledge AcquisitionXinran Gu, Kaifeng Lyu, Jiazheng Li, Jingzhao ZhangNeurIPS 2025 · 被引用 16 次
它引用的顶会 Paper23
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 被引用 3,415 次
- Transformers Learn In-Context by Gradient DescentJohannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento 等ICML 2023 · 被引用 729 次
- Hopfield Networks is All You NeedHubert Ramsauer, Bernhard Schäfl, Johannes Lehner, Philipp Seidl 等ICLR 2021 · 被引用 620 次
- Linear Transformers Are Secretly Fast Weight ProgrammersImanol Schlag, Kazuki Irie, Jürgen SchmidhuberICML 2021 · 被引用 394 次
- Transformers learn to implement preconditioned gradient descent for in-context learningKwangjun Ahn, Xiang Cheng, Hadi Daneshmand, Suvrit SraNeurIPS 2023 · 被引用 324 次
相关 Paper
- Do LLMs dream of elephants (when told not to)? Latent concept association and associative memory in transformersYibo Jiang, Goutham Rajendran, Pradeep Ravikumar, Bryon AragamNeurIPS 2024 · 被引用 19 次
- Learning to Recall with Transformers Beyond Orthogonal EmbeddingsNuri Mert Vural, Alberto Bietti, Mahdi Soltanolkotabi, Denny WuICLR 2026 · 被引用 1 次
- Memory Layers at ScaleVincent-Pierre Berges, Barlas Oguz, Daniel Haziza, Wen-tau Yih 等ICML 2025
- Do All Autoregressive Transformers Remember Facts the Same Way? A Cross-Architecture Analysis of Recall MechanismsMinyeong Choe, Haehyun Cho, Changho Seo, Hyunil KimEMNLP 2025
- Birth of a Transformer: A Memory ViewpointAlberto Bietti, Vivien Cabannes, Diane Bouchacourt, Hervé Jégou 等NeurIPS 2023 · 被引用 182 次
