AdaSplash: Adaptive Sparse Flash Attention
Nuno Gonçalves, Marcos V. Treviso, André F. T. Martins
摘要
The computational cost of softmax-based attention in transformers limits their applicability to long-context tasks. Adaptive sparsity, of which α-entmax attention is an example, offers a flexible data-dependent alternative, but existing implementations are inefficient and do not leverage the sparsity to obtain runtime and memory gains. In this work, we propose ADASPLASH, which combines the efficiency of GPU-optimized algorithms with the sparsity benefits of α-entmax. We first introduce a hybrid Halley-bisection algorithm, resulting in a 7-fold reduction in the number of iterations needed to compute the α-entmax transformation. Then, we implement custom Triton kernels to efficiently handle adaptive sparsity. Experiments with RoBERTa and ModernBERT for text classification and single-vector retrieval, along with GPT-2 for language modeling, show that our method achieves substantial improvements in runtime and memory efficiency compared to existing α-entmax implementations. It approachesand in some cases surpasses-the efficiency of highly optimized softmax implementations like FlashAttention-2, enabling long-context training while maintaining strong task performance. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Long-Context Generalization with Sparse AttentionPavlo Vasylenko, Hugo Pitorro, Andre F. T. Martins, Marcos V. TrevisoICLR 2026 · 被引用 19 次
- Threshold Differential Attention for Sink-Free, Ultra-Sparse, and Non-Dispersive Language ModelingXingyue Huang, Xueying Ding, Mingxuan Ju, Yozen Liu 等ACL 2026 · 被引用 3 次
- AdaSplash-2: Faster Differentiable Sparse AttentionNuno M. T. Gonçalves, Hugo Pitorro, Vlad Niculae, Edoardo Ponti 等ICML 2026 · 被引用 3 次
- Improving Sparse Autoencoder with Dynamic AttentionDongsheng Wang, Jinsen Zhang, Dawei Su, Hui HuangCVPR 2026 · 被引用 2 次
- SSA: Sparse Sparse Attention by Aligning Full and Sparse Attention Outputs in Feature SpaceZhenyi Shen, Junru Lu, Lin Gui, Jiazheng Li 等ICML 2026 · 被引用 2 次
它引用的顶会 Paper17
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie 等NeurIPS 2020 · 被引用 3,159 次
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 被引用 2,878 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
相关 Paper
- SALO: an efficient spatial accelerator enabling hybrid sparse attention mechanisms for long sequencesGuan Shen, Jieru Zhao, Quan Chen, Jingwen Leng 等DAC 2022 · 被引用 37 次
- FSA: An Alternative Efficient Implementation of Native Sparse Attention KernelRan Yan, Youhe Jiang, Zhuoming Chen, Haohui Mai 等ICLR 2026 · 被引用 10 次
- Fast Attention Over Long Sequences With Dynamic Sparse Flash AttentionMatteo Pagliardini, Daniele Paliotta, Martin Jaggi, François FleuretNeurIPS 2023 · 被引用 26 次
- Scaling Attention via Feature SparsityYan Xie, Tiansheng Wen, Tangda Huang, Bo Chen 等ICLR 2026 · 被引用 3 次
- A Unified Sparse Attention via Multi-Granularity CompressionSiran Liu, Zheng Cao, Yongchao HeICML 2026
