Accelerating Retrieval Augmented Language Model via PIM and PNM Integration
Je-Woo Jang, Junyong Oh, Youngbae Kong, Jae-Youn Hong, Sung-Hyuk Cho, Jeongyeol Lee, Hoeseok Yang, Joon-Sung Yang
Abstract
Retrieval-Augmented Language Models (RALMs) integrate a language model with an external database to generate high-quality outputs utilizing up-to-date information.However, both components of a RALM system, the language model and the retriever, suffer from distinct memory-bound bottlenecks.In particular, the attention mechanism of the language model heavily relies on General Matrix-Vector Multiplication (GEMV) operations using unique K/V matrices per request, complicating batch parallelization and exacerbating memory bandwidth constraints.Conversely, the retriever encounters performance bottlenecks due to frequent LUT lookups and intensive sorting operations, characterized by low arithmetic intensity and limited data reuse, making GPU acceleration challenging.To address these distinctive characteristics, this paper proposes MNM, a hardware architecture integrating Processing In Memory (PIM) within the HBM core die and Processing Near Memory (PNM) on the HBM logic die.The PIM module leverages the high internal bandwidth of HBM to accelerate GEMV operations in the language model, while the PNM module optimizes retrieval-specific tasks.Furthermore, this work introduces a novel RALM scheduling strategy combining selective batching and early generation to exploit the performance improvements achieved by the MNM architecture.By strategically overlapping retrieval and generation phases, the proposed scheduling scheme reduces idle cycles in a batched RALM system.Experimental results demonstrate that the proposed techniques achieve up to 29.2× performance speedup compared to a conventional GPU-based RALM system.In addition, the proposed PIM/PNM-integrated approach saves up to 71.5% of energy consumption, highlighting its applicability for memory-bound RALM workloads.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get bdb27b10-b8c6-46e0-84fb-ec26ba74c3c7Cited by top-tier papers1
Ask how each one uses itRelated papers
- Move Less, Retrieve Fast: A Retrieval-in-Memory Architecture for Language ModelsJiaxian Chen, Yuxuan Qi, Jianan Yuan, Kaoyi Sun et al.DAC 2025 · 2 citations
- HeterRAG: Heterogeneous Processing-in-Memory Acceleration for Retrieval-augmented GenerationChaoqiang Liu, Haifeng Liu, Dan Chen, Yu Huang et al.ISCA 2025 · 10 citations
- Chameleon: a Heterogeneous and Disaggregated Accelerator System for Retrieval-Augmented Language ModelsWenqi Jiang, Marco Zeller, Roger Waleffe, Torsten Hoefler et al.VLDB 2025 · 50 citations
- In-Storage Acceleration of Retrieval Augmented Generation as a ServiceRohan Mahapatra, Harsha Santhanam, Christopher Priebe, Hanyang Xu et al.ISCA 2025 · 9 citations
- DReX: Accurate and Scalable Dense Retrieval Acceleration via Algorithmic-Hardware CodesignDerrick Quinn, E. Ezgi Yücel, Martin Prammer, Zhenxing Fan et al.ISCA 2025 · 10 citations
