Lune

DAC2021Top-tier venue

MAT: Processing In-Memory Acceleration for Long-Sequence Attention

Minxuan Zhou, Yunhui Guo, Weihong Xu, Bin Li, Kevin W. Eliceiri, Tajana Rosing

2021Year
10Citations
1Top-tier citations

Abstract

Attention-based machine learning is used to model long-term dependencies in sequential data. Processing these models on long sequences can be prohibitively costly because of the large memory consumption. In this work, we propose MAT, a processing in-memory (PIM) framework, to accelerate long-sequence attention models. MAT adopts a memory-efficient processing flow for attention models to process sub-sequences in a pipeline with much smaller memory footprint. MAT utilizes a reuse-driven data layout and an optimal sample scheduling to optimize the performance of PIM attention. We evaluate the efficiency of MAT on two emerging long-sequence tasks including natural language processing and medical image processing. Our experiments show that MAT is 2.7×2.7 \times faster and 3.4×3.4 \times more energy efficient than the state-of-the-art PIM acceleration. As compared to TPU and GPU, MAT is 5.1×5.1 \times and 16.4×16.4 \times faster while consuming 27.5×27.5 \times and 41.0×41.0 \times less energy.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

Cited by top-tier papers1

Ask how each one uses it

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines