MAT: Processing In-Memory Acceleration for Long-Sequence Attention
Minxuan Zhou, Yunhui Guo, Weihong Xu, Bin Li, Kevin W. Eliceiri, Tajana Rosing
Abstract
Attention-based machine learning is used to model long-term dependencies in sequential data. Processing these models on long sequences can be prohibitively costly because of the large memory consumption. In this work, we propose MAT, a processing in-memory (PIM) framework, to accelerate long-sequence attention models. MAT adopts a memory-efficient processing flow for attention models to process sub-sequences in a pipeline with much smaller memory footprint. MAT utilizes a reuse-driven data layout and an optimal sample scheduling to optimize the performance of PIM attention. We evaluate the efficiency of MAT on two emerging long-sequence tasks including natural language processing and medical image processing. Our experiments show that MAT is faster and more energy efficient than the state-of-the-art PIM acceleration. As compared to TPU and GPU, MAT is and faster while consuming and less energy.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers1
Ask how each one uses itRelated papers
- FLAT: An Optimized Dataflow for Mitigating Attention BottlenecksSheng-Chun Kao, Suvinay Subramanian, Gaurav Agrawal, Amir Yazdanbakhsh et al.ASPLOS 2023 · 68 citations
- BlockPIM: Optimizing Memory Management for PIM-enabled Long-Context LLM InferenceZhichun Li, Jun Zhou, Xueqi Li, Ninghui SunDAC 2025 · 3 citations
- PAISE: PIM-Accelerated Inference Scheduling Engine for Transformer-based LLMHyojung Lee, Daehyeon Baek, Jimyoung Son, Jieun Choi et al.HPCA 2025 · 10 citations
- Accelerating Sparse Attention with a Reconfigurable Non-volatile Processing-In-Memory ArchitectureQilin Zheng, Shiyu Li, Yitu Wang, Ziru Li et al.DAC 2023 · 14 citations
- STARC: Selective Token Access with Remapping and Clustering for Efficient LLM Decoding on PIM SystemsZehao Fan, Yunzhen Liu, Garrett Gagnon, Zhenyu Liu et al.ASPLOS 2026
