Meridian: In-Memory Acceleration for RAG with Document Attention Decomposition
Chaoqiang Liu, Yu Huang, Haifeng Liu, Yi Zhang, Qihang Qiu, Xueqi Li, Long Zheng, Xiaofei Liao, Hai Jin, Jingling Xue
Abstract
Retrieval-Augmented Generation (RAG) improves the factuality and timeliness of large language model outputs by incorporating external knowledge during inference. Recent systems accelerate RAG by precomputing and caching documentside Key-Value (KV) pairs, eliminating repeated encoding of long retrieved documents. However, this centralized KV-reuse paradigm introduces two fundamental bottlenecks: (1) massive off-chip KV transfers, since large-scale document KVs must reside in host memory and be moved to the device at query time, and (2) severely underutilized compute resources, as short queries yield skinny GEMMs during prefilling and memory-bound GEMVs during decoding. We address these limitations with Meridian, a decentralized RAG system built on two key components. First, we introduce document attention decomposition, which replaces centralized KV processing with a distributed execution model: document-side and matrices are sharded across PIM-enabled memory modules, and each device computes attention over its local shard, producing compact partial summaries that are merged through a lightweight global aggregation step. This sharply reduces offchip KV movement. Second, to improve compute efficiency, Meridian incorporates a PIM-based accelerator co-designed with the decomposition mechanism. It provides a resource-conscious in-memory compute substrate for accelerating skinny GEMM and nonlinear operations, and employs a coordination-aware hybrid scheduler to sustain efficient intra-device execution and scalable inter-device parallelism. Evaluations show that Meridian achieves average throughput improvements of 5.36×/6.64×/3.98×/3.32×/3.91× and latency reductions of 4.30×/5.34×/3.31×/2.73×/2.79× over TurboRAG, BlockAttention, CENT, PAPI, and HeterRAG, respectively.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- HeterRAG: Heterogeneous Processing-in-Memory Acceleration for Retrieval-augmented GenerationChaoqiang Liu, Haifeng Liu, Dan Chen, Yu Huang et al.ISCA 2025 · 10 citations
- In-Storage Acceleration of Retrieval Augmented Generation as a ServiceRohan Mahapatra, Harsha Santhanam, Christopher Priebe, Hanyang Xu et al.ISCA 2025 · 9 citations
- LazyAttention: Efficient Retrieval-Augmented Generation with Deferred Positional EncodingHaocheng Xia, Mihir Pamnani, Hanxi Fang, Supawit Chockchowwat et al.ICML 2026
- DReX: Accurate and Scalable Dense Retrieval Acceleration via Algorithmic-Hardware CodesignDerrick Quinn, E. Ezgi Yücel, Martin Prammer, Zhenxing Fan et al.ISCA 2025 · 10 citations
- TurboRAG: Accelerating Retrieval-Augmented Generation with Precomputed KV Caches for Chunked TextSongshuo Lu, Hua Wang, Yutian Rong, Zhi Chen et al.EMNLP 2025 · 2 citations
