AttAcc! Unleashing the Power of PIM for Batched Transformer-based Generative Model Inference
Jaehyun Park, Jaewan Choi, Kwanhee Kyung, Michael Jaemin Kim, Yongsuk Kwon, Nam Sung Kim, Jung Ho Ahn
Abstract
The Transformer-based generative model (TbGM), comprising summarization (Sum) and generation (Gen) stages, has demonstrated unprecedented generative performance across a wide range of applications. However, it also demands immense amounts of compute and memory resources. Especially, the Gen stages, consisting of the attention and fully-connected (FC) layers, dominate the overall execution time. Meanwhile, we reveal that the conventional system with GPUs used for TbGM inference cannot efficiently execute the attention layer, even with batching, due to various constraints. To address this inefficiency, we first propose AttAcc, a processing-in-memory (PIM) architecture for efficient execution of the attention layer. Subsequently, for the end-to-end acceleration of TbGM inference, we propose a novel heterogeneous system architecture and optimizations that strategically use xPU and PIM together. It leverages the high memory bandwidth of AttAcc for the attention layer and the powerful compute capability of the conventional system for the FC layer. Lastly, we demonstrate that our GPU-PIM system outperforms the conventional system with the same memory capacity, improving performance and energy efficiency of running a 175B TbGM by up to 2.81× and 2.67×, respectively.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers34
- PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model InferenceYufeng Gu, Alireza Khadem, Sumanth Umesh, Ning Liang et al.ASPLOS 2025 · 44 citations
- IMPRESS: An Importance-Informed Multi-Tier Prefix KV Storage System for Large Language Model InferenceWeijian Chen, Shuibing He, Haoyang Qu, Ruidong Zhang et al.FAST 2025 · 40 citations
- Duplex: A Device for Large Language Models with Mixture of Experts, Grouped Query Attention, and Continuous BatchingSungmin Yun, Kwanhee Kyung, Juhwan Cho, Jaewan Choi et al.MICRO 2024 · 40 citations
- PAPI: Exploiting Dynamic Parallelism in Large Language Model Decoding with a Processing-In-Memory-Enabled Computing SystemYintao He, Haiyu Mao, Christina Giannoula, Mohammad Sadrosadati et al.ASPLOS 2025 · 37 citations
- Oaken: Fast and Efficient LLM Serving with Online-Offline Hybrid KV Cache QuantizationMinsu Kim, Seongmin Hong, Ryeowook Ko, Soongyu Choi et al.ISCA 2025 · 17 citations
Related papers
- PAISE: PIM-Accelerated Inference Scheduling Engine for Transformer-based LLMHyojung Lee, Daehyeon Baek, Jimyoung Son, Jieun Choi et al.HPCA 2025 · 10 citations
- IANUS: Integrated Accelerator based on NPU-PIM Unified Memory SystemMinseok Seo, Xuan Truong Nguyen, Seok Joong Hwang, Yongkee Kwon et al.ASPLOS 2024 · 57 citations
- DFX: A Low-latency Multi-FPGA Appliance for Accelerating Transformer-based Text GenerationSeongmin Hong, Seungjae Moon, Junsoo Kim, Sungjae Lee et al.MICRO 2022 · 107 citations
- TransPIM: A Memory-based Acceleration via Software-Hardware Co-Design for TransformerMinxuan Zhou, Weihong Xu, Jaeyoung Kang, Tajana RosingHPCA 2022 · 142 citations
- NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM InferencingGuseul Heo, Sangyeop Lee, Jaehong Cho, Hyunmin Choi et al.ASPLOS 2024 · 121 citations
