Lune

DAC2021顶会

MAT: Processing In-Memory Acceleration for Long-Sequence Attention

Minxuan Zhou, Yunhui Guo, Weihong Xu, Bin Li, Kevin W. Eliceiri, Tajana Rosing

2021年份
10被引次数
1顶会引用

摘要

Attention-based machine learning is used to model long-term dependencies in sequential data. Processing these models on long sequences can be prohibitively costly because of the large memory consumption. In this work, we propose MAT, a processing in-memory (PIM) framework, to accelerate long-sequence attention models. MAT adopts a memory-efficient processing flow for attention models to process sub-sequences in a pipeline with much smaller memory footprint. MAT utilizes a reuse-driven data layout and an optimal sample scheduling to optimize the performance of PIM attention. We evaluate the efficiency of MAT on two emerging long-sequence tasks including natural language processing and medical image processing. Our experiments show that MAT is 2.7×2.7 \times faster and 3.4×3.4 \times more energy efficient than the state-of-the-art PIM acceleration. As compared to TPU and GPU, MAT is 5.1×5.1 \times and 16.4×16.4 \times faster while consuming 27.5×27.5 \times and 41.0×41.0 \times less energy.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper1

问问它们各自怎么用它

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖