COSA:Co-Operative Systolic Arrays for Multi-head Attention Mechanism in Neural Network using Hybrid Data Reuse and Fusion Methodologies
Zhican Wang, Gang Wang, Honglan Jiang, Ningyi Xu, Guanghui He
Abstract
Attention mechanism acceleration is becoming increasingly vital to achieve superior performance in deep learning tasks. Existing accelerators are commonly devised dedicatedly by exploring the potential sparsity in neural network (NN) models, which suffer from complicated training, tuning processes, and accuracy degradation. By systematically analyzing the inherent dataflow characteristics of attention mechanism, we propose the Co-Operative Systolic Array (COSA) to pursue higher computational efficiency for its acceleration. In COSA, two systolic arrays that can be dynamically configured into weight or output stationary modes are cascaded to enable efficient attention operation. Thus, hybrid dataflows are simultaneously supported in COSA. Furthermore, various fusion methodologies and an advanced softmax unit are designed. Experimental results show that the COSA-based accelerator can achieve 2.95-28.82× speedup compared with the existing designs, with up to 97.4% PE utilization rate and less memory access.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers1
Ask how each one uses itRelated papers
- An Energy-Efficient High-Utilization Hardware Architecture for Attention Mechanism in Transformer using Balanced Systolic Array and Multi-Row Interleaved Operation OrderingHaiyang Zhou, Hongyang Hu, Jinshan Yue, Hanghang Gao et al.DAC 2025 · 1 citation
- Blaze: An Efficient Bit-Sparse Attention Architecture With Workload Orchestration OptimizationRunzhou Zhang, Faxian Sun, Yiming Wang, Kunchen Zou et al.DAC 2025
- DEFA: Efficient Deformable Attention Acceleration via Pruning-Assisted Grid-Sampling and Multi-Scale Parallel ProcessingYansong Xu, Dongxu Lyu, Zhenyu Li, Yuzhou Chen et al.DAC 2024 · 5 citations
- Libra: A Hybrid-Sparse Attention Accelerator Featuring Multi-Level Workload BalanceFaxian Sun, Runzhou Zhang, Zhenyu Liu, Heng Liao et al.DAC 2025 · 1 citation
- SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head PruningHanrui Wang, Zhekai Zhang, Song HanHPCA 2021 · 412 citations
