∞-former: Infinite Memory Transformer
Pedro Henrique Martins, Zita Marinho, André F. T. Martins
Abstract
Transformers are unable to model long-term memories effectively, since the amount of computation they need to perform grows with the context length. While variations of efficient transformers have been proposed, they all have a finite memory capacity and are forced to drop old information. In this paper, we propose the ∞-former, which extends the vanilla transformer with an unbounded longterm memory. By making use of a continuousspace attention mechanism to attend over the long-term memory, the ∞-former's attention complexity becomes independent of the context length, trading off memory length with precision. In order to control where precision is more important, ∞-former maintains "sticky memories," being able to model arbitrarily long contexts while keeping the computation budget fixed. Experiments on a synthetic sorting task, language modeling, and document grounded dialogue generation demonstrate the ∞-former's ability to retain information from long sequences. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9d04bf93-8fad-4d11-a4bf-893b7ba08e68Cited by top-tier papers17
- Recurrent Memory TransformerAydar Bulatov, Yuri Kuratov, Mikhail BurtsevNeurIPS 2022 · 252 citations
- MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory AgentHongli Yu, Tinghong Chen, Jiangtao Feng, Jiangjie Chen et al.ICLR 2026 · 231 citations
- LongBench: A Bilingual, Multitask Benchmark for Long Context UnderstandingYushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu et al.ACL 2024 · 94 citations
- Memory Consolidation Enables Long-Context Video UnderstandingIvana Balazevic, Yuge Shi, Pinelopi Papalampidi, Rahma Chaabouni et al.ICML 2024 · 52 citations
- Training Language Models with Memory AugmentationZexuan Zhong, Tao Lei, Danqi ChenEMNLP 2022 · 52 citations
Builds on12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
- Generalization through Memorization: Nearest Neighbor Language ModelsUrvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer et al.ICLR 2020 · 1,038 citations
- Compressive Transformers for Long-Range Sequence ModellingJack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, Chloe Hillier et al.ICLR 2020 · 833 citations
Related papers
- ∞-Video: A Training-Free Approach to Long Video Understanding via Continuous-Time Memory ConsolidationSaul José Rodrigues dos Santos, António Farinhas, Daniel C. McNamee, André F. T. MartinsICML 2025
- Unlimiformer: Long-Range Transformers with Unlimited Length InputAmanda Bertsch, Uri Alon, Graham Neubig, Matthew GormleyNeurIPS 2023 · 176 citations
- Nyströmformer: A Nyström-based Algorithm for Approximating Self-AttentionYunyang Xiong, Zhanpeng Zeng, Rudrasis Chakraborty, Mingxing Tan et al.AAAI 2021 · 675 citations
- Proxyformer: Nyström-Based Linear Transformer with Trainable Proxy TokensSangho Lee, Hayun Lee, Dongkun ShinAAAI 2024 · 5 citations
- Every Token Counts: Generalizing 16M Ultra-Long Context in Large Language ModelsXiang Hu, Zhanchao Zhou, Ruiqi Liang, Zehuan Li et al.ACL 2026 · 2 citations
