Efficient Length-Generalizable Attention via Causal Retrieval for Long-Context Language Modeling
Xiang Hu, Zhihao Teng, Jun Zhao, Wei Wu, Kewei Tu
Abstract
Despite the success of Transformers, handling longer contexts remains challenging due to the limited length generalization and quadratic complexity of self-attention. Transformers often require post-training with a larger attention window, significantly increasing computational and memory costs. In this paper, we propose a novel attention mechanism based on dynamic context, Grouped Cross Attention (GCA), which can generalize to 1000 × the pre-training context length while maintaining the ability to access distant information with a constant attention window size. For a given input sequence, we split it into chunks and use each chunk to retrieve top-k relevant past chunks for subsequent text generation. Specifically, unlike most previous works that use an offthe-shelf retriever, our key innovation allows the retriever to learn how to retrieve past chunks that better minimize the auto-regressive loss of subsequent tokens in an end-to-end manner. Such a mechanism accommodates retrieved chunks with a fixed-size attention window to achieve longrange information access, significantly reducing computational and memory costs during training and inference. Experiments show that GCA-based models achieve near-perfect accuracy in passkey retrieval for 16M context lengths, which is 1000× the training length.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Long-Context Generalization with Sparse AttentionPavlo Vasylenko, Hugo Pitorro, Andre F. T. Martins, Marcos V. TrevisoICLR 2026 · 19 citations
- Hardware-aligned Hierarchical Sparse Attention for Efficient Long-term Memory AccessXiang Hu, Jiaqi Leng, Jun Zhao, Kewei Tu et al.NeurIPS 2025 · 7 citations
- Understanding and Improving Length Generalization in Hierarchical Sparse Attention ModelsJiaqi Leng, Xiang Hu, Junxiong Wang, Jianguo Li et al.ICLR 2026 · 6 citations
- Every Token Counts: Generalizing 16M Ultra-Long Context in Large Language ModelsXiang Hu, Zhanchao Zhou, Ruiqi Liang, Zehuan Li et al.ACL 2026 · 2 citations
Builds on20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
- Improving Language Models by Retrieving from Trillions of TokensSebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai et al.ICML 2022 · 1,629 citations
Related papers
- Reducing Transformer Key-Value Cache Size with Cross-Layer AttentionWilliam Brandon, Mayank Mishra, Aniruddha Nrusimha, Rameswar Panda et al.NeurIPS 2024 · 150 citations
- Random-Access Infinite Context Length for TransformersAmirkeivan Mohtashami, Martin JaggiNeurIPS 2023 · 207 citations
- Core Context Aware Transformers for Long Context Language ModelingYaofo Chen, Zeng You, Shuhai Zhang, Haokun Li et al.ICML 2025
- Training-Free Long-Context Scaling of Large Language ModelsChenxin An, Fei Huang, Jun Zhang, Shansan Gong et al.ICML 2024 · 68 citations
- Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language ModelingLiliang Ren, Yang Liu, Yadong Lu, Yelong Shen et al.ICLR 2025
