Long-Context Generalization with Sparse Attention
Pavlo Vasylenko, Hugo Pitorro, Andre F. T. Martins, Marcos V. Treviso
摘要
Transformer-based architectures traditionally employ softmax to compute attention weights, which produces dense distributions over all tokens in a sequence. While effective in many settings, this density has been shown to be detrimental for tasks that demand precise focus on fixed-size patterns: as sequence length increases, noninformative tokens accumulate attention probability mass, leading to dispersion and representational collapse. We show in this paper that dynamically sparse attention mechanisms using α-entmax can avoid these issues, due to their ability to assign exact zeros to irrelevant tokens. Furthermore, we introduce Adaptive-Scalable Entmax (ASEntmax), which endows α-entmax with a learnable temperature parameter, allowing the attention distribution to interpolate between sparse (pattern-focused) and dense (softmax-like) regimes. Our empirical evaluation on synthetic tasks and language modeling demonstrates that ASEntmax substantially outperforms softmax, scalable softmax, and fixed-temperature α-entmax baselines, achieving up to 1000× length extrapolation on synthetic benchmarks and superior long-context generalization on language modeling while preserving short-context performance, including better perplexity trends and higher retrieval accuracies at 8× training length. Source code: https://github.com/deep-spin/asentmax
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- TabICLv2: A Better, Faster, Scalable, and Open Tabular Foundation ModelJingang QU, David Holzmüller, Gael Varoquaux, Marine Le MorvanICML 2026 · 被引用 85 次
- VideoNSA: Native Sparse Attention Scales Video UnderstandingEnxin Song, Wenhao Chai, Shusheng Yang, Ethan Armand 等ICLR 2026 · 被引用 11 次
- AdaSplash-2: Faster Differentiable Sparse AttentionNuno M. T. Gonçalves, Hugo Pitorro, Vlad Niculae, Edoardo Ponti 等ICML 2026 · 被引用 3 次
- Modelling Attention with Aitchison Geometry: Token Distinguishability and Temperature ScalingSam Hilton-Jones, Timothy Norman, Zhanxing ZhuICML 2026
它引用的顶会 Paper39
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie 等NeurIPS 2020 · 被引用 3,159 次
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
相关 Paper
- Sparse and Continuous Attention MechanismsAndré F. T. Martins, António Farinhas, Marcos V. Treviso, Vlad Niculae 等NeurIPS 2020 · 被引用 55 次
- Sparse Text GenerationPedro Henrique Martins, Zita Marinho, André F. T. MartinsEMNLP 2020 · 被引用 20 次
- AdaSplash: Adaptive Sparse Flash AttentionNuno Gonçalves, Marcos V. Treviso, André F. T. MartinsICML 2025
- Efficient Transformer Inference with Statically Structured Sparse AttentionSteve Dai, Hasan Genc, Rangharajan Venkatesan, Brucek KhailanyDAC 2023 · 被引用 11 次
- Selective Attention: Enhancing Transformer through Principled Context ControlXuechen Zhang, Xiangyu Chang, Mingchen Li, Amit K. Roy-Chowdhury 等NeurIPS 2024 · 被引用 32 次
