Beyond Attention: Breaking the Limits of Transformer Context Length with Recurrent Memory
Aydar Bulatov, Yuri Kuratov, Yermek Kapushev, Mikhail Burtsev
Abstract
A major limitation for the broader scope of problems solvable by transformers is the quadratic scaling of computational complexity with input size. In this study, we investigate the recurrent memory augmentation of pre-trained transformer models to extend input context length while linearly scaling compute. Our approach demonstrates the capability to store information in memory for sequences of up to an unprecedented two million tokens while maintaining high retrieval accuracy. Experiments with language modeling tasks show perplexity improvement as the number of processed input segments increases. These results underscore the effectiveness of our method, which has significant potential to enhance long-term dependency handling in natural language understanding and generation tasks, as well as enable large-scale context processing for memory-intensive applications.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 310e4131-517a-4c1c-9524-c625e8a3af1aCited by top-tier papers3
- Scale-invariant attentionBen Anson, Xi Wang, Laurence AitchisonNeurIPS 2025 · 6 citations
- VideoTitans: Scalable Video Prediction with Integrated Short- and Long-term MemoryYoung-Jae Park, Minseok Seo, Hae-Gon JeonNeurIPS 2025 · 4 citations
- InfiniPot: Infinite Context Processing on Memory-Constrained LLMsMinsoo Kim, Kyuhong Shim, Jungwook Choi, Simyung ChangEMNLP 2024 · 1 citation
Builds on9
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- Compressive Transformers for Long-Range Sequence ModellingJack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, Chloe Hillier et al.ICLR 2020 · 833 citations
- An empirical analysis of compute-optimal large language model trainingJordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya et al.NeurIPS 2022 · 566 citations
- Recurrent Memory TransformerAydar Bulatov, Yuri Kuratov, Mikhail BurtsevNeurIPS 2022 · 252 citations
- Unlimiformer: Long-Range Transformers with Unlimited Length InputAmanda Bertsch, Uri Alon, Graham Neubig, Matthew GormleyNeurIPS 2023 · 176 citations
Related papers
- Random-Access Infinite Context Length for TransformersAmirkeivan Mohtashami, Martin JaggiNeurIPS 2023 · 207 citations
- Block-Recurrent TransformersDeLesley Hutchins, Imanol Schlag, Yuhuai Wu, Ethan Dyer et al.NeurIPS 2022 · 163 citations
- Improving Language Models by Retrieving from Trillions of TokensSebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai et al.ICML 2022 · 1,629 citations
- Pretraining with hierarchical memories: separating long-tail and common knowledgeHadi Pouransari, David Grangier, C Thomas, Michael Kirchhof et al.ICLR 2026 · 11 citations
- Augmenting Language Models with Long-Term MemoryWeizhi Wang, Li Dong, Hao Cheng, Xiaodong Liu et al.NeurIPS 2023 · 256 citations
