Move Less, Retrieve Fast: A Retrieval-in-Memory Architecture for Language Models
Jiaxian Chen, Yuxuan Qi, Jianan Yuan, Kaoyi Sun, Tianyu Wang, Chenlin Ma, Yi Wang
Abstract
Retrieval-augmented language models (RALMs) have attracted widespread attention for addressing the limitations of traditional large language models. However, challenges involved in retrieval, including substantial data movement and irregular access patterns, seriously impact the efficiency and deployment of RALMs. The emerging 3D-stacked processing-in-memory (PIM) architecture, characterized by its high memory bandwidth and near-data computing capabilities, presents a promising solution for efficient retrieval. To support large-scale retrieval in RALMs, the PIM architecture should be carefully designed with joint software and hardware optimization. This paper presents Rimast, a retrieval-in-memory architecture for fast retrieval in RALMs. The objective is to minimize data movement and improve overall performance through hardwaresoftware co-design. At the hardware level, a hierarchical PIM architecture with a retrieval-in-memory dataflow is designed to reduce unnecessary data transfer. At the software level, skew-free data mapping and adaptive offloading strategies are proposed to address the irregular access patterns associated with retrieval in RALMs. We demonstrate the effectiveness of the proposed Rimast using extensive experiments. The experimental results demonstrate that Rimast effectively reduces data movement, achieving average speedups of , and over CPUs, GPUs, and prior art accelerators, respectively.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 4b605853-4a53-48e0-adbc-339586e9abe0Related papers
- Accelerating Retrieval Augmented Language Model via PIM and PNM IntegrationJe-Woo Jang, Junyong Oh, Youngbae Kong, Jae-Youn Hong et al.MICRO 2025 · 5 citations
- DReX: Accurate and Scalable Dense Retrieval Acceleration via Algorithmic-Hardware CodesignDerrick Quinn, E. Ezgi Yücel, Martin Prammer, Zhenxing Fan et al.ISCA 2025 · 10 citations
- Meridian: In-Memory Acceleration for RAG with Document Attention DecompositionChaoqiang Liu, Yu Huang, Haifeng Liu, Yi Zhang et al.ISCA 2026 · 2 citations
- HeterRAG: Heterogeneous Processing-in-Memory Acceleration for Retrieval-augmented GenerationChaoqiang Liu, Haifeng Liu, Dan Chen, Yu Huang et al.ISCA 2025 · 10 citations
- Accelerating Iterative Retrieval-augmented Language Model Serving with SpeculationZhihao Zhang, Alan Zhu, Lijie Yang, Yihua Xu et al.ICML 2024 · 10 citations
