AttenPIM: Accelerating LLM Attention with Dual-mode GEMV in Processing-in-Memory
Liyan Chen, Dongxu Lyu, Zhenyu Li, Jianfei Jiang, Qin Wang, Zhigang Mao, Naifeng Jing
Abstract
Large Language Models (LLMs) have demonstrated unprecedented generative performance across a wide range of applications. While recent heterogeneous architectures attempt to address the memory-bound bottleneck from attention computations by processing-in-memory (PIM) offloading, they overlook two critical characteristics of attention GEMVs that distinguish them from traditional PIM scenarios: (1) dynamic matrix dimensions that scale with token length, and (2) distinct GEMV patterns between score computation () and context computation (). Existing PIM designs, employing either uniform or transposed computing modes, suffer from inefficiencies in newly generated element preparation or distinct GEMV execution. To address these limitations, we propose AttenPIM, a software-hardware co-design for efficient PIM-based attention acceleration. For bank-level execution, we propose dual-mode computing modes tailored for score and context computations with PIM-oriented data layouts and execution flows for KV storage, supported by a low-cost configurable per-bank PIM unit (PU). For system-level execution, we leverage token-level and head-level concurrency to ensure workload balance and maximize bank PU parallelism. Furthermore, dynamic allocation and kernel fusion methods are proposed to further minimize memory overhead. Experimental results demonstrate that AttenPIM achieves speedup and reduces energy consumption by 17 %-49 % compared to two state-of-the-art PIM baselines.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 815c435c-309c-4c7f-a548-d3e8125cb28bCited by top-tier papers1
Ask how each one uses itRelated papers
- BlockPIM: Optimizing Memory Management for PIM-enabled Long-Context LLM InferenceZhichun Li, Jun Zhou, Xueqi Li, Ninghui SunDAC 2025 · 3 citations
- NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM InferencingGuseul Heo, Sangyeop Lee, Jaehong Cho, Hyunmin Choi et al.ASPLOS 2024 · 121 citations
- PAISE: PIM-Accelerated Inference Scheduling Engine for Transformer-based LLMHyojung Lee, Daehyeon Baek, Jimyoung Son, Jieun Choi et al.HPCA 2025 · 10 citations
- Bridging Efficiency and Scalability in Llm System Via 3D Hybrid Pim With 2D in-Transit ComputationHongyi Li, Songchen Ma, Huanyu Qu, Weihao Zhang et al.ISCA 2026
- AQPIM: Breaking the PIM Capacity Wall for LLMs with in-Memory Activation QuantizationKosuke Matsushima, Yasuyuki Okoshi, Masato Motomura, Daichi FujikiHPCA 2026 · 1 citation
