Lune

MICRO2022顶会

Sparse Attention Acceleration with Synergistic In-Memory Pruning and On-Chip Recomputation

Amir Yazdanbakhsh, Ashkan Moradifirouzabadi, Zheng Li, Mingu Kang

2022年份
47被引次数
17顶会引用

摘要

comparisons against threshold values can outweigh the benefits of in-memory computing.

  1. Selective read of unpruned embeddings: Supporting inmemory ReRAM pruning enforces a particular data layout for key embeddings. However, this layout constraints the ability to selectively read the unpruned vectors. To remedy these considerations, this work makes the following contributions: 1 We introduce a unique perspective on the ReRAM in-memory computing paradigm. We employ approximate in-memory compute and precise on-chip recompute in tandem to mitigate the likely negative repercussions to model accuracy due to inherent circuit inaccuracies. 2 We employ analog comparators to carry out the comparisons with threshold values and instead produce 1-bit data to indicate the pruning status. With this shift in design, we reduce the hardware cost, which is proportional to input bit precision, to merely the cost of a series of 1-bit analog to digital converters (ADCs). 3 We repurpose an existing solution, which enables us to implement data reuse based on our observations. On the hardware side, we rely on recently taped-out transposable ReRAMs [141] that introduce in-situ transposed read access. While initially intended for efficiently accessing neural network weights, our application of this hardware selectively reads unpruned embeddings. For the data reuse, we observe that there is a considerable spatial locality between unpruned key vectors of adjacent queries. We exploit this spatial location to improve data reuse and further reduce the data communication overhead.

We evaluate our approach in several self-attention models with large sequences, including BERT, ALBERT, ViT, GPT-2, and two futuristic designs (e.g. 2K and 4K input sequence length). Under an iso design, our results show that, on average, SPRINT delivers 7.5× speed-up and 19.6× energy reduction compared to a baseline design with 16KB on-chip memory. The benefit increases as on-chip resources become scarcer, representing a design point for resource constrained platforms, e.g. 1.6× more energy reduction with 16KB on-chip memory than the case with 64KB capacity.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext 38b08e1f-2549-4a4d-8570-2da94d64042e

引用它的顶会 Paper17

问问它们各自怎么用它

它引用的顶会 Paper26

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖