PAISE: PIM-Accelerated Inference Scheduling Engine for Transformer-based LLM
Hyojung Lee, Daehyeon Baek, Jimyoung Son, Jieun Choi, Kihyo Moon, Minsung Jang
Abstract
Transformer-based Large Language Models (LLMs) demand significant computational and memory resources due to the autoregressive token generation in decoder blocks. In particular, the attention layer in LLM models has low arithmetic intensity but high memory traffic, thus requiring frequent updates to the KV matrices with each decoder iteration. As a result, LLM inference becomes memory bound, leading to increased latency. To address this, we introduce PAISE, a framework leveraging Processing-In-Memory (PIM) technology to offload memory-intensive tasks. PAISE employs GPU-PIM heterogeneous computing resources to optimize inference operations in transformer-based LLMs. The framework comprises (i) a scheduling algorithm that decides which operations to offload to PIM based on model configuration and PIM hardware specifications and (ii) an enhanced PIM kernel that performs transaction-wise interleave-batched GEMM (General Matrix Multiplication) operations, maximizing data throughput via data layout adjustments. We implemented PAISE on the GPT-2 and Llama2-7B models using an AMD MI100 GPU with HBM-PIM devices. Our evaluations show that offloading the attention layer to PIM reduces execution time by up to 48.3% compared to GPU-only inference, demonstrating PAISE’s significant potential to enhance the efficiency of LLM inference, which could lead to faster and more efficient AI applications.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 1e9c689c-c5f7-404d-9288-634eacaa62deCited by top-tier papers5
- CHIME: A Case for Efficient Long-Context Attention-FC Disaggregated Inference with DIMM-PIMQingyuan Liu, Liyan Chen, Haocheng Wang, Yanning Yang et al.ISCA 2026 · 6 citations
- DCC: Data-Centric Compilation of Machine Learning Kernels for Processing-In-Memory ArchitecturesPeiming Yang, Sankeerth Durvasula, Ivan Fernandez, Mohammad Sadrosadati et al.ISCA 2026 · 3 citations
- Cocoon: A System Architecture for Differentially Private Training with Correlated NoisesDonghwan Kim, Xin Gu, Jinho Baek, Timothy Lo et al.OSDI 2026 · 1 citation
- PIM-Malloc: A Fast and Scalable Dynamic Memory Allocator for Processing-In-Memory (PIM) ArchitecturesDongjae Lee, Bongjoon Hyun, Youngjin Kwon, Minsoo RhuHPCA 2026 · 1 citation
- A Cost-Effective Near-Storage Processing Solution for Offline Inference of Long-Context LLMsHongsun Jang, Jaeyong Song, Changmin Shin, Si Ung Noh et al.ASPLOS 2026
Related papers
- PAPI: Exploiting Dynamic Parallelism in Large Language Model Decoding with a Processing-In-Memory-Enabled Computing SystemYintao He, Haiyu Mao, Christina Giannoula, Mohammad Sadrosadati et al.ASPLOS 2025 · 37 citations
- NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM InferencingGuseul Heo, Sangyeop Lee, Jaehong Cho, Hyunmin Choi et al.ASPLOS 2024 · 121 citations
- Understand and Accelerate Memory Processing Pipeline for Large Language Model InferenceZifan He, Rui Ma, Yizhou Sun, Jason CongICML 2026
- AttAcc! Unleashing the Power of PIM for Batched Transformer-based Generative Model InferenceJaehyun Park, Jaewan Choi, Kwanhee Kyung, Michael Jaemin Kim et al.ASPLOS 2024 · 125 citations
- IANUS: Integrated Accelerator based on NPU-PIM Unified Memory SystemMinseok Seo, Xuan Truong Nguyen, Seok Joong Hwang, Yongkee Kwon et al.ASPLOS 2024 · 57 citations
