ASADI: Accelerating Sparse Attention Using Diagonal-based In-Situ Computing
Huize Li, Zhaoying Li, Zhenyu Bai, Tulika Mitra
摘要
The self-attention mechanism is the performance bottleneck of Transformer-based language models, particularly for long sequences. Researchers have proposed using sparse attention to speed up the Transformer. However, sparse attention introduces significant random access overhead, limiting computational efficiency. To mitigate this issue, researchers attempt to improve data reuse by utilizing row/column locality. Unfortunately, we find that sparse attention does not naturally exhibit strong row/column locality, but instead has excellent diagonal locality. Thus, it is worthwhile to use diagonal compression (DIA) format. However, existing sparse matrix computation paradigms struggle to efficiently support DIA format in attention computation. To address this problem, we propose ASADI, a novel software-hardware co-designed sparse attention accelerator. In the soft-ware side, we propose a new sparse matrix computation paradigm that directly supports the DIA format in self-attention computation. In the hardware side, we present a novel sparse attention accelerator that efficiently implements our computation paradigm using highly parallel in-situ computing. We thoroughly evaluate ASADI across various models and datasets. Our experimental results demonstrate an average performance improvement of 18.6 × and energy savings of 2.9× compared to a PIM-based baseline.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper11
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie 等NeurIPS 2020 · 被引用 3,159 次
- You Only Look at One Sequence: Rethinking Transformer in Vision through Object DetectionYuxin Fang, Bencheng Liao, Xinggang Wang, Jiemin Fang 等NeurIPS 2021 · 被引用 430 次
- A3: Accelerating Attention Mechanisms in Neural Networks with ApproximationTae Jun Ham, Sungjun Jung, Seonghak Kim, Young H. Oh 等HPCA 2020 · 被引用 241 次
- Sanger: A Co-Design Framework for Enabling Sparse Attention using Reconfigurable ArchitectureLiqiang Lu, Yicheng Jin, Hangrui Bi, Zizhang Luo 等MICRO 2021 · 被引用 221 次
相关 Paper
- SpARC: Token Similarity-Aware Sparse Attention Transformer Accelerator via Row-wise ClusteringHan Cho, Dongjun Kim, Seung-Eon Hwang, Jongsun ParkDAC 2024 · 被引用 7 次
- SOFA: A Compute-Memory Optimized Sparsity Accelerator via Cross-Stage Coordinated TilingHuizheng Wang, Jiahao Fang, Xinru Tang, Zhiheng Yue 等MICRO 2024 · 被引用 31 次
- An Energy-Efficient High-Utilization Hardware Architecture for Attention Mechanism in Transformer using Balanced Systolic Array and Multi-Row Interleaved Operation OrderingHaiyang Zhou, Hongyang Hu, Jinshan Yue, Hanghang Gao 等DAC 2025 · 被引用 1 次
- A length adaptive algorithm-hardware co-design of transformer on FPGA through sparse attention and dynamic pipeliningHongwu Peng, Shaoyi Huang, Shiyang Chen, Bingbing Li 等DAC 2022 · 被引用 49 次
- SALO: an efficient spatial accelerator enabling hybrid sparse attention mechanisms for long sequencesGuan Shen, Jieru Zhao, Quan Chen, Jingwen Leng 等DAC 2022 · 被引用 37 次
