TSAcc: An Efficient Tempo-Spatial Similarity Aware Accelerator for Attention Acceleration
Zhuoran Song, Chunyu Qi, Yuanzheng Yao, Peng Zhou, Yanyi Zi, Nan Wang, Xiaoyao Liang
摘要
Attention-based models provide significant accuracy improvement to Natural Language Processing (NLP) and computer vision (CV) fields at the cost of heavy computational and memory demands. Previous works seek to alleviate the performance bottleneck by removing useless relations for each position. However, their attempts only focus on intra-sentence optimization and overlook the opportunity in the temporal domain. In this paper, we accelerate attention by leveraging the tempo-spatial similarity across successive sentences, given the observation that successive sentences tend to bear high similarity. This is rational owing to many semantic similar words (namely tokens) in the attention-based models. We first propose an online-offline prediction algorithm to identify similar tokens/heads. We then design a recovery algorithm so that we can skip the computation on similar tokens/heads in succeeding sentences and recover their results by copying other tokens/heads features in preceding sentences to reserve accuracy. From the hardware aspect, we propose a specialized architecture TSAcc that includes a prediction engine and recovery engine to translate the computational saving in the algorithm to real speedup. Experiments show that TSAcc can achieve 8.5X, 2.7X, 14.1X, and 64.9X speedup compared to SpAtten, Sanger, 1080TI GPU, and Xeon CPU, with negligible accuracy loss.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- CTA: Hardware-Software Co-design for Compressed Token Attention MechanismHaoran Wang, Haobo Xu, Ying Wang, Yinhe HanHPCA 2023 · 被引用 18 次
- A Memory-Efficient LLM Accelerator with Q-K Correlation Prediction using Cluster-Based Associative Array for Selective KV AccessingZikang Zhou, Kaiqi Chen, Xuyang Duan, Jun HanDAC 2025 · 被引用 2 次
- SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head PruningHanrui Wang, Zhekai Zhang, Song HanHPCA 2021 · 被引用 412 次
- ELSA: Hardware-Software Co-design for Efficient, Lightweight Self-Attention Mechanism in Neural NetworksTae Jun Ham, Yejin Lee, Seong Hoon Seo, Soosung Kim 等ISCA 2021 · 被引用 185 次
- ASADI: Accelerating Sparse Attention Using Diagonal-based In-Situ ComputingHuize Li, Zhaoying Li, Zhenyu Bai, Tulika MitraHPCA 2024 · 被引用 22 次
