An Energy-Efficient High-Utilization Hardware Architecture for Attention Mechanism in Transformer using Balanced Systolic Array and Multi-Row Interleaved Operation Ordering
Haiyang Zhou, Hongyang Hu, Jinshan Yue, Hanghang Gao, Yuanlu Xie, Xiaoxin Xu, Chunmeng Dou, Ming Liu
摘要
Transformer-based neural networks have achieved remarkable performance. Designing energy-efficient and high-speed accelerators for the attention mechanism, which dominates the energy and latency in Transformers, has become increasingly significant. Existing attention accelerators commonly use algorithm-hardware co-design to achieve higher energy efficiency and speed. However, deeply customized algorithms make these accelerators dependent on a particular application. Therefore, optimizing hardware architecture is crucial for achieving general-purpose acceleration. We observe two limitations in the hardware architecture of existing attention accelerators. First, the widely used input stationary, weight stationary, and output stationary systolic arrays (SAs) can’t balance data reuse, register saving, and utilization, which hinders to build more energy-efficient and faster SA-based accelerators. Second, layer-by-layer operation ordering introduces high SRAM access overhead of intermediate results. To address the first limitation, we propose the “Balanced Systolic Array”, which improves energy efficiency by 40% compared to conventional systolic arrays and achieves a utilization rate of 99.5%. To address the second limitation, we propose “Multi-Row Interleaved” operation ordering, which reduces the SRAM energy by 31.7% By integrating two techniques, the proposed attention accelerator achieves a 39% improvement in energy efficiency and a 38% enhancement in throughput×energy efficiency compared to previous works.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- COSA:Co-Operative Systolic Arrays for Multi-head Attention Mechanism in Neural Network using Hybrid Data Reuse and Fusion MethodologiesZhican Wang, Gang Wang, Honglan Jiang, Ningyi Xu 等DAC 2023 · 被引用 14 次
- FACT: FFN-Attention Co-optimized Transformer Architecture with Eager Correlation PredictionYubin Qin, Yang Wang, Dazheng Deng, Zhiren Zhao 等ISCA 2023 · 被引用 113 次
- A length adaptive algorithm-hardware co-design of transformer on FPGA through sparse attention and dynamic pipeliningHongwu Peng, Shaoyi Huang, Shiyang Chen, Bingbing Li 等DAC 2022 · 被引用 49 次
- SpARC: Token Similarity-Aware Sparse Attention Transformer Accelerator via Row-wise ClusteringHan Cho, Dongjun Kim, Seung-Eon Hwang, Jongsun ParkDAC 2024 · 被引用 7 次
- ASADI: Accelerating Sparse Attention Using Diagonal-based In-Situ ComputingHuize Li, Zhaoying Li, Zhenyu Bai, Tulika MitraHPCA 2024 · 被引用 22 次
