SPLAT: A Framework for Optimised GPU Code-Generation for SParse reguLar ATtention
Ahan Gupta, Yueming Yuan, Devansh Jain, Yuhao Ge, David Aponte, Yanqi Zhou, Charith Mendis
摘要
Multi-head-self-attention (MHSA) mechanisms achieve state-of-the-art (SOTA) performance across natural language processing and vision tasks. However, their quadratic dependence on sequence lengths has bottlenecked inference speeds. To circumvent this bottleneck, researchers have proposed various sparse-MHSA models, where a subset of full attention is computed. Despite their promise, current sparse libraries and compilers do not support high-performance implementations for diverse sparse-MHSA patterns due to the underlying sparse formats they operate on. These formats, which are typically designed for high-performance & scientific computing applications, are either curated for extreme amounts of random sparsity (<1% non-zero values), or specific sparsity patterns. However, the sparsity patterns in sparse-MHSA are moderately sparse (10-50% non-zero values) and varied, resulting in existing sparse-formats trading off generality for performance.
We bridge this gap, achieving both generality and performance, by proposing a novel sparse format: affine-compressed-sparse-row (ACSR) and supporting code-generation scheme, SPLAT, that generates highperformance implementations for diverse sparse-MHSA patterns on GPUs. Core to our proposed format and code generation algorithm is the observation that common sparse-MHSA patterns have uniquely regular geometric properties. These properties, which can be analyzed just-in-time, expose novel optimizations and tiling strategies that SPLAT exploits to generate high-performance implementations for diverse patterns. To demonstrate SPLAT's efficacy, we use it to generate code for various sparse-MHSA models, achieving geomean speedups of 2.05x and 4.05x over hand-written kernels written in Triton and TVM respectively on A100 GPUs. Moreover, its interfaces are intuitive and easy to use with existing implementations of MHSA in JAX.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- DCC: Data-Centric Compilation of Machine Learning Kernels for Processing-In-Memory ArchitecturesPeiming Yang, Sankeerth Durvasula, Ivan Fernandez, Mohammad Sadrosadati 等ISCA 2026 · 被引用 3 次
- Accelerating Sparse Transformer Inference on GPUWenhao Dai, Haodong Deng, Mengfei Rong, Xinyu Yang 等PPoPP 2026 · 被引用 1 次
- HASTE: Hardware-Aware Dynamic Sparse Training for Large Output SpacesNasib Ullah, Jinbin Zhang, Jean Lucien Randrianantenaina, Erik Schultheis 等ICML 2026
- Characterizing Real-World Bugs in Tile Programs for Automated Bug DetectionRavishka Rathnasuriya, Zihe Song, Nidhi Majoju, Tingxi Li 等ISSTA 2026
- SAS: Sparse Attention Synthesizer for Efficient Language Model InferenceYuan Zhou, Shaojie Xiang, Lingfan Yu, Zhenyu Song 等EuroSys 2026
它引用的顶会 Paper10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 被引用 2,878 次
- Long Range Arena : A Benchmark for Efficient TransformersYi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen 等ICLR 2021 · 被引用 881 次
- LongLoRA: Efficient Fine-tuning of Long-Context Large Language ModelsYukang Chen, Shengju Qian, Haotian Tang, Xin Lai 等ICLR 2024 · 被引用 254 次
- Long-Short Transformer: Efficient Transformers for Language and VisionChen Zhu, Wei Ping, Chaowei Xiao, Mohammad Shoeybi 等NeurIPS 2021 · 被引用 180 次
相关 Paper
- FSA: An Alternative Efficient Implementation of Native Sparse Attention KernelRan Yan, Youhe Jiang, Zhuoming Chen, Haohui Mai 等ICLR 2026 · 被引用 10 次
- VENOM: A Vectorized N: M Format for Unleashing the Power of Sparse Tensor CoresRoberto L. Castro, Andrei Ivanov, Diego Andrade, Tal Ben-Nun 等SC 2023 · 被引用 22 次
- ASADI: Accelerating Sparse Attention Using Diagonal-based In-Situ ComputingHuize Li, Zhaoying Li, Zhenyu Bai, Tulika MitraHPCA 2024 · 被引用 22 次
- SpargeAttention: Accurate and Training-free Sparse Attention Accelerating Any Model InferenceJintao Zhang, Chendong Xiang, Haofeng Huang, Jia Wei 等ICML 2025
- SparseD: Sparse Attention for Diffusion Language ModelsZeqing Wang, Gongfan Fang, Xinyin Ma, Xingyi Yang 等ICLR 2026 · 被引用 18 次
