REPA: Reconfigurable PIM for the Joint Acceleration of KV Cache Offloading and Processing
Yang Hong, Junlong Yang, Bo Peng, Jianguo Yao
摘要
The use of KV cache in LLM inference leads to large memory footprint and sub-optimal decoding performance. Prior studies typically address one of these two limitations by either offloading or stage-split inference. In this paper, we explore and reveal the possibility of a joint solution, and propose REPA, a GPU-PIM hybrid system to prototype this idea. We leverage reconfigurable ReRAM PIM to achieve fast KV cache persistence, and balance the requirement of processing speed and memory capacity. To fully unleash the parallelization potential of REPA, we propose optimizations in (1) architecture, (2) data mapping and (3) pipelining: (1) We propose bulk-wise memory instructions and multi-level controllers to enable finer-grained parallelism in the PIM device. (2) We propose locality-aware data mapping to make the best of the aforementioned architectural optimization, and reduce long-range data transfer on chip. (3) We adopt sub-batch pipelining to reduce idleness in batches, and propose transfer overlapping to shadow the KV cache transfer by computation. Experimental results show that REPA exhibits high inference speed, energy efficiency and integratability. It is 1.5--6.5× faster, and 8--10× more efficient than NVIDIA A100. It also outperforms state-of-the-art DRAM PIM systems by up to 1.4× for long context inference. When integrated into existing offloading systems, REPA achieves 1.4--2.0× offloading speed, and 1.2--1.4× end-to-end speedup, showcasing its high potential for fast KV cache offloading and processing.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- KVDrive: A Holistic Multi-Tier KV Cache Management System for Long-Context LLM InferenceJian Lin, Jiazhi Mi, Zicong Hong, Haodong Wang 等SIGMOD 2026 · 被引用 5 次
- PKAS: Predictive KVCache-Aware Scheduling for Faster LLM and Transformer InferencesJie Ye, Avinash Maurya, Krishna Teja Chitty-Venkata, Bogdan Nicolae 等HPDC 2026
- No Buffer, No Bottleneck: Efficient Zero-Copy KV Cache Offloading for Long-Context LLMsShutian Luo, Haiying ShenOSDI 2026
- Scaling Attention Beyond GPUs for LLM InferenceWeishu Deng, Yujie Yang, Peiran Du, Lingfeng Xiang 等HPDC 2026
- Efficient Remote KV Cache Reuse with GPU-native Video CodecLiang Mi, Weijun Wang, Jinghan Chen, Ting Cao 等SIGCOMM 2026
