Meridian: In-Memory Acceleration for RAG with Document Attention Decomposition
Chaoqiang Liu, Yu Huang, Haifeng Liu, Yi Zhang, Qihang Qiu, Xueqi Li, Long Zheng, Xiaofei Liao, Hai Jin, Jingling Xue
摘要
Retrieval-Augmented Generation (RAG) improves the factuality and timeliness of large language model outputs by incorporating external knowledge during inference. Recent systems accelerate RAG by precomputing and caching documentside Key-Value (KV) pairs, eliminating repeated encoding of long retrieved documents. However, this centralized KV-reuse paradigm introduces two fundamental bottlenecks: (1) massive off-chip KV transfers, since large-scale document KVs must reside in host memory and be moved to the device at query time, and (2) severely underutilized compute resources, as short queries yield skinny GEMMs during prefilling and memory-bound GEMVs during decoding. We address these limitations with Meridian, a decentralized RAG system built on two key components. First, we introduce document attention decomposition, which replaces centralized KV processing with a distributed execution model: document-side and matrices are sharded across PIM-enabled memory modules, and each device computes attention over its local shard, producing compact partial summaries that are merged through a lightweight global aggregation step. This sharply reduces offchip KV movement. Second, to improve compute efficiency, Meridian incorporates a PIM-based accelerator co-designed with the decomposition mechanism. It provides a resource-conscious in-memory compute substrate for accelerating skinny GEMM and nonlinear operations, and employs a coordination-aware hybrid scheduler to sustain efficient intra-device execution and scalable inter-device parallelism. Evaluations show that Meridian achieves average throughput improvements of 5.36×/6.64×/3.98×/3.32×/3.91× and latency reductions of 4.30×/5.34×/3.31×/2.73×/2.79× over TurboRAG, BlockAttention, CENT, PAPI, and HeterRAG, respectively.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- HeterRAG: Heterogeneous Processing-in-Memory Acceleration for Retrieval-augmented GenerationChaoqiang Liu, Haifeng Liu, Dan Chen, Yu Huang 等ISCA 2025 · 被引用 10 次
- In-Storage Acceleration of Retrieval Augmented Generation as a ServiceRohan Mahapatra, Harsha Santhanam, Christopher Priebe, Hanyang Xu 等ISCA 2025 · 被引用 9 次
- LazyAttention: Efficient Retrieval-Augmented Generation with Deferred Positional EncodingHaocheng Xia, Mihir Pamnani, Hanxi Fang, Supawit Chockchowwat 等ICML 2026
- DReX: Accurate and Scalable Dense Retrieval Acceleration via Algorithmic-Hardware CodesignDerrick Quinn, E. Ezgi Yücel, Martin Prammer, Zhenxing Fan 等ISCA 2025 · 被引用 10 次
- TurboRAG: Accelerating Retrieval-Augmented Generation with Precomputed KV Caches for Chunked TextSongshuo Lu, Hua Wang, Yutian Rong, Zhi Chen 等EMNLP 2025 · 被引用 2 次
