Efficient Length-Generalizable Attention via Causal Retrieval for Long-Context Language Modeling
Xiang Hu, Zhihao Teng, Jun Zhao, Wei Wu, Kewei Tu
摘要
Despite the success of Transformers, handling longer contexts remains challenging due to the limited length generalization and quadratic complexity of self-attention. Transformers often require post-training with a larger attention window, significantly increasing computational and memory costs. In this paper, we propose a novel attention mechanism based on dynamic context, Grouped Cross Attention (GCA), which can generalize to 1000 × the pre-training context length while maintaining the ability to access distant information with a constant attention window size. For a given input sequence, we split it into chunks and use each chunk to retrieve top-k relevant past chunks for subsequent text generation. Specifically, unlike most previous works that use an offthe-shelf retriever, our key innovation allows the retriever to learn how to retrieve past chunks that better minimize the auto-regressive loss of subsequent tokens in an end-to-end manner. Such a mechanism accommodates retrieved chunks with a fixed-size attention window to achieve longrange information access, significantly reducing computational and memory costs during training and inference. Experiments show that GCA-based models achieve near-perfect accuracy in passkey retrieval for 16M context lengths, which is 1000× the training length.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Long-Context Generalization with Sparse AttentionPavlo Vasylenko, Hugo Pitorro, Andre F. T. Martins, Marcos V. TrevisoICLR 2026 · 被引用 19 次
- Hardware-aligned Hierarchical Sparse Attention for Efficient Long-term Memory AccessXiang Hu, Jiaqi Leng, Jun Zhao, Kewei Tu 等NeurIPS 2025 · 被引用 7 次
- Understanding and Improving Length Generalization in Hierarchical Sparse Attention ModelsJiaqi Leng, Xiang Hu, Junxiong Wang, Jianguo Li 等ICLR 2026 · 被引用 6 次
- Every Token Counts: Generalizing 16M Ultra-Long Context in Large Language ModelsXiang Hu, Zhanchao Zhou, Ruiqi Liang, Zehuan Li 等ACL 2026 · 被引用 2 次
它引用的顶会 Paper20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
- Improving Language Models by Retrieving from Trillions of TokensSebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai 等ICML 2022 · 被引用 1,629 次
相关 Paper
- Reducing Transformer Key-Value Cache Size with Cross-Layer AttentionWilliam Brandon, Mayank Mishra, Aniruddha Nrusimha, Rameswar Panda 等NeurIPS 2024 · 被引用 150 次
- Random-Access Infinite Context Length for TransformersAmirkeivan Mohtashami, Martin JaggiNeurIPS 2023 · 被引用 207 次
- Core Context Aware Transformers for Long Context Language ModelingYaofo Chen, Zeng You, Shuhai Zhang, Haokun Li 等ICML 2025
- Training-Free Long-Context Scaling of Large Language ModelsChenxin An, Fei Huang, Jun Zhang, Shansan Gong 等ICML 2024 · 被引用 68 次
- Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language ModelingLiliang Ren, Yang Liu, Yadong Lu, Yelong Shen 等ICLR 2025
