Long-Context Generalization with Sparse Attention
Pavlo Vasylenko, Hugo Pitorro, Andre F. T. Martins, Marcos V. Treviso
Abstract
Transformer-based architectures traditionally employ softmax to compute attention weights, which produces dense distributions over all tokens in a sequence. While effective in many settings, this density has been shown to be detrimental for tasks that demand precise focus on fixed-size patterns: as sequence length increases, noninformative tokens accumulate attention probability mass, leading to dispersion and representational collapse. We show in this paper that dynamically sparse attention mechanisms using α-entmax can avoid these issues, due to their ability to assign exact zeros to irrelevant tokens. Furthermore, we introduce Adaptive-Scalable Entmax (ASEntmax), which endows α-entmax with a learnable temperature parameter, allowing the attention distribution to interpolate between sparse (pattern-focused) and dense (softmax-like) regimes. Our empirical evaluation on synthetic tasks and language modeling demonstrates that ASEntmax substantially outperforms softmax, scalable softmax, and fixed-temperature α-entmax baselines, achieving up to 1000× length extrapolation on synthetic benchmarks and superior long-context generalization on language modeling while preserving short-context performance, including better perplexity trends and higher retrieval accuracies at 8× training length. Source code: https://github.com/deep-spin/asentmax
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fa6fafe1-e0f7-4c0e-949a-030a9b36a9dbCited by top-tier papers4
- TabICLv2: A Better, Faster, Scalable, and Open Tabular Foundation ModelJingang QU, David Holzmüller, Gael Varoquaux, Marine Le MorvanICML 2026 · 85 citations
- VideoNSA: Native Sparse Attention Scales Video UnderstandingEnxin Song, Wenhao Chai, Shusheng Yang, Ethan Armand et al.ICLR 2026 · 11 citations
- AdaSplash-2: Faster Differentiable Sparse AttentionNuno M. T. Gonçalves, Hugo Pitorro, Vlad Niculae, Edoardo Ponti et al.ICML 2026 · 3 citations
- Modelling Attention with Aitchison Geometry: Token Distinguishability and Temperature ScalingSam Hilton-Jones, Timothy Norman, Zhanxing ZhuICML 2026
Builds on39
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
Related papers
- Sparse and Continuous Attention MechanismsAndré F. T. Martins, António Farinhas, Marcos V. Treviso, Vlad Niculae et al.NeurIPS 2020 · 55 citations
- Sparse Text GenerationPedro Henrique Martins, Zita Marinho, André F. T. MartinsEMNLP 2020 · 20 citations
- AdaSplash: Adaptive Sparse Flash AttentionNuno Gonçalves, Marcos V. Treviso, André F. T. MartinsICML 2025
- Efficient Transformer Inference with Statically Structured Sparse AttentionSteve Dai, Hasan Genc, Rangharajan Venkatesan, Brucek KhailanyDAC 2023 · 11 citations
- Selective Attention: Enhancing Transformer through Principled Context ControlXuechen Zhang, Xiangyu Chang, Mingchen Li, Amit K. Roy-Chowdhury et al.NeurIPS 2024 · 32 citations
