Random-Access Infinite Context Length for Transformers
Amirkeivan Mohtashami, Martin Jaggi
Abstract
While Transformers have shown remarkable success in natural language processing, their attention mechanism's large memory requirements have limited their ability to handle longer contexts. Prior approaches, such as recurrent memory or retrieval-based augmentation, have either compromised the random-access flexibility of attention (i.e., the capability to select any token in the entire context) or relied on separate mechanisms for relevant context retrieval, which may not be compatible with the model's attention. In this paper, we present a novel approach that allows access to the complete context while retaining random-access flexibility, closely resembling running attention on the entire context. Our method uses a landmark token to represent each block of the input and trains the attention to use it for selecting relevant blocks, enabling retrieval of blocks directly through the attention mechanism instead of by relying on a separate mechanism. Our approach seamlessly integrates with specialized data structures and the system's memory hierarchy, enabling processing of arbitrarily long context lengths. We demonstrate that our method can obtain comparable performance with Transformer-XL while significantly reducing the number of retrieved tokens in each step. Finally, we show that fine-tuning LLaMA 7B with our method successfully extends its context length capacity to over 32k tokens, allowing for inference at the context lengths of GPT-4. We release the implementation of landmark attention and the code to reproduce our experiments at https://github.com/epfml/landmark-attention/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5df22200-ea0d-4e93-905c-f61e4ac83ef1Cited by top-tier papers54
- Neural Attention SearchDifan Deng, Marius LindauerNeurIPS 2025 · 431 citations
- Make Your LLM Fully Utilize the ContextShengnan An, Zexiong Ma, Zeqi Lin, Nanning Zheng et al.NeurIPS 2024 · 212 citations
- Focused Transformer: Contrastive Training for Context ScalingSzymon Tworkowski, Konrad Staniszewski, Mikolaj Pacek, Yuhuai Wu et al.NeurIPS 2023 · 190 citations
- Reducing Transformer Key-Value Cache Size with Cross-Layer AttentionWilliam Brandon, Mayank Mishra, Aniruddha Nrusimha, Rameswar Panda et al.NeurIPS 2024 · 150 citations
- Base of RoPE Bounds Context LengthMingyu Xu, Xin Men, Bingning Wang, Qingyu Zhang et al.NeurIPS 2024 · 56 citations
Builds on15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- Train Short, Test Long: Attention with Linear Biases Enables Input Length ExtrapolationOfir Press, Noah A. Smith, Mike LewisICLR 2022 · 1,168 citations
Related papers
- Training-Free Long-Context Scaling of Large Language ModelsChenxin An, Fei Huang, Jun Zhang, Shansan Gong et al.ICML 2024 · 68 citations
- ReAttention: Training-Free Infinite Context with Finite Attention ScopeXiaoran Liu, Ruixiao Li, Zhigeng Liu, Qipeng Guo et al.ICLR 2025
- Landmark Embedding: A Chunking-Free Embedding Method For Retrieval Augmented Long-Context Large Language ModelsKun Luo, Zheng Liu, Shitao Xiao, Tong Zhou et al.ACL 2024 · 10 citations
- Beyond Attention: Breaking the Limits of Transformer Context Length with Recurrent MemoryAydar Bulatov, Yuri Kuratov, Yermek Kapushev, Mikhail BurtsevAAAI 2024 · 22 citations
- Efficient Length-Generalizable Attention via Causal Retrieval for Long-Context Language ModelingXiang Hu, Zhihao Teng, Jun Zhao, Wei Wu et al.ICML 2025
