Lune

MICRO2022Top-tier venue

Sparse Attention Acceleration with Synergistic In-Memory Pruning and On-Chip Recomputation

Amir Yazdanbakhsh, Ashkan Moradifirouzabadi, Zheng Li, Mingu Kang

2022Year
47Citations
17Top-tier citations

Abstract

comparisons against threshold values can outweigh the benefits of in-memory computing.

  1. Selective read of unpruned embeddings: Supporting inmemory ReRAM pruning enforces a particular data layout for key embeddings. However, this layout constraints the ability to selectively read the unpruned vectors. To remedy these considerations, this work makes the following contributions: 1 We introduce a unique perspective on the ReRAM in-memory computing paradigm. We employ approximate in-memory compute and precise on-chip recompute in tandem to mitigate the likely negative repercussions to model accuracy due to inherent circuit inaccuracies. 2 We employ analog comparators to carry out the comparisons with threshold values and instead produce 1-bit data to indicate the pruning status. With this shift in design, we reduce the hardware cost, which is proportional to input bit precision, to merely the cost of a series of 1-bit analog to digital converters (ADCs). 3 We repurpose an existing solution, which enables us to implement data reuse based on our observations. On the hardware side, we rely on recently taped-out transposable ReRAMs [141] that introduce in-situ transposed read access. While initially intended for efficiently accessing neural network weights, our application of this hardware selectively reads unpruned embeddings. For the data reuse, we observe that there is a considerable spatial locality between unpruned key vectors of adjacent queries. We exploit this spatial location to improve data reuse and further reduce the data communication overhead.

We evaluate our approach in several self-attention models with large sequences, including BERT, ALBERT, ViT, GPT-2, and two futuristic designs (e.g. 2K and 4K input sequence length). Under an iso design, our results show that, on average, SPRINT delivers 7.5× speed-up and 19.6× energy reduction compared to a baseline design with 16KB on-chip memory. The benefit increases as on-chip resources become scarcer, representing a design point for resource constrained platforms, e.g. 1.6× more energy reduction with 16KB on-chip memory than the case with 64KB capacity.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 38b08e1f-2549-4a4d-8570-2da94d64042e

Cited by top-tier papers17

Ask how each one uses it

Builds on26

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines