DPU-KV: On the Benefits of DPU Offloading for In-Memory Key-Value Stores at the Edge
Arjun Kashyap, Yuke Li, Xiaoyi Lu
摘要
In-memory key-value stores (KVS) are widely used for edge data storage, where low latency and high throughput are essential. Data Processing Units (DPUs), with their low power use and offloading capabilities, suit resource-constrained edge computing. While DPUs offer a new design point for KVS, their offloading in edge environments remains underexplored and challenging. In this paper, we unveil the potential of offloading in-memory CPU-based KVS to SoC-based DPUs, specifically NVIDIA's BlueField-2 (BF-2) and BlueField-3 (BF-3), with the aim of enhancing KVS performance. We propose a principled exploration methodology of dividing a KVS (i.e., MICA) into its logical components and identifying the CPU-intensive KVS component (i.e., communication engine). Next, we perform fine-grained offloading analysis and explorations on DPUs. To maximize benefits in terms of latency and throughput from fine-grained KVS offloading on DPUs, we propose a series of significant performance optimizations, including a key-value-based queue-pair model, overlapped KV request/response processing, reduced DMA operations per KV batch, dual-communication engine, and a sharding-based design. Our key finding is that our proposed fine-grained KVS offloading designs on modern DPU architectures (i.e., BF-2 and BF-3) can provide much lower latency (up to 68%) and higher throughput (up to 36%) than MICA (CPU-only) and coarse-grained DPU offloading schemes at the edge. To our knowledge, this paper is the first to explore the performance benefits of fine-grained KVS offloading to DPUs at the edge.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- Accelerating Sketch-based End-Host Traffic Measurement with Automatic DPU OffloadingXiang Chen, Xi Sun, Wenbin Zhang, Xin Yao 等INFOCOM 2024 · 被引用 5 次
- DDS: DPU-optimized Disaggregated StorageQizhen Zhang, Philip A. Bernstein, Badrish Chandramouli, Jason Hu 等VLDB 2024 · 被引用 12 次
- PD3: Prefetching Data with DPUs for Disaggregated MemorySidharth Sankhe, Felix Zhang, Umayrah Chonee, Sherman Lim 等NSDI 2026 · 被引用 1 次
- Characterizing Off-path SmartNIC for Accelerating Distributed SystemsXingda Wei, Rongxin Cheng, Yuhan Yang, Rong Chen 等OSDI 2023 · 被引用 68 次
- Analyzing Near-Network Hardware Acceleration with Co-Processing on DPUsDimitrios Giouroukis, Dwi P. A. Nugroho, Varun Pandey, Steffen Zeuch 等VLDB 2025 · 被引用 3 次
