Move Less, Retrieve Fast: A Retrieval-in-Memory Architecture for Language Models
Jiaxian Chen, Yuxuan Qi, Jianan Yuan, Kaoyi Sun, Tianyu Wang, Chenlin Ma, Yi Wang
摘要
Retrieval-augmented language models (RALMs) have attracted widespread attention for addressing the limitations of traditional large language models. However, challenges involved in retrieval, including substantial data movement and irregular access patterns, seriously impact the efficiency and deployment of RALMs. The emerging 3D-stacked processing-in-memory (PIM) architecture, characterized by its high memory bandwidth and near-data computing capabilities, presents a promising solution for efficient retrieval. To support large-scale retrieval in RALMs, the PIM architecture should be carefully designed with joint software and hardware optimization. This paper presents Rimast, a retrieval-in-memory architecture for fast retrieval in RALMs. The objective is to minimize data movement and improve overall performance through hardwaresoftware co-design. At the hardware level, a hierarchical PIM architecture with a retrieval-in-memory dataflow is designed to reduce unnecessary data transfer. At the software level, skew-free data mapping and adaptive offloading strategies are proposed to address the irregular access patterns associated with retrieval in RALMs. We demonstrate the effectiveness of the proposed Rimast using extensive experiments. The experimental results demonstrate that Rimast effectively reduces data movement, achieving average speedups of , and over CPUs, GPUs, and prior art accelerators, respectively.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Accelerating Retrieval Augmented Language Model via PIM and PNM IntegrationJe-Woo Jang, Junyong Oh, Youngbae Kong, Jae-Youn Hong 等MICRO 2025 · 被引用 5 次
- DReX: Accurate and Scalable Dense Retrieval Acceleration via Algorithmic-Hardware CodesignDerrick Quinn, E. Ezgi Yücel, Martin Prammer, Zhenxing Fan 等ISCA 2025 · 被引用 10 次
- Meridian: In-Memory Acceleration for RAG with Document Attention DecompositionChaoqiang Liu, Yu Huang, Haifeng Liu, Yi Zhang 等ISCA 2026 · 被引用 2 次
- HeterRAG: Heterogeneous Processing-in-Memory Acceleration for Retrieval-augmented GenerationChaoqiang Liu, Haifeng Liu, Dan Chen, Yu Huang 等ISCA 2025 · 被引用 10 次
- Accelerating Iterative Retrieval-augmented Language Model Serving with SpeculationZhihao Zhang, Alan Zhu, Lijie Yang, Yihua Xu 等ICML 2024 · 被引用 10 次
