DFlush: DPU-Offloaded Flush for Disaggregated LSM-based Key-Value Stores
Chen Ding, Kai Lu, Quanyi Zhang, Zekun Ye, Ting Yao, Daohui Wang, Huatao Wu, Jiguang Wan
Abstract
Rapid increase of storage and network bandwidth incurs higher CPU consumption in modern data systems. This phenomenon is particularly evident for log-structured merged key-value stores (LSM-KVS), which rely on resource-intensive background operations to flush and compact disk data. While extensive research has been conducted to reduce the CPU overhead of background compaction, less attention has been paid to background flushing, which can also consume a significant amount of valuable CPU cycles and disrupt CPU caches, ultimately impacting overall performance. In this paper, we propose DFlush, a novel solution that uses DPUs to offload background flush operations to reduce its CPU cost. DPUs are an appealing choice for this goal due to their cost-effectiveness, ease of programming, and widespread deployment. However, their complex hardware architecture requires careful design of both the data and control planes. To fully harness the DPU's capabilities, DFlush decomposes a flush job into fine-grained steps, mapped them to DPU hardware units, and accelerates them through pipeline, data, and channel parallelism, ensuring data-plane efficiency. It also introduces an adaptive control plane that dynamically schedules flush jobs from different LSM-KVS instances based on their priority, reducing write stall and tail latency. Our experiments on a real DPU platform with an industrial-grade LSM-KVS show that DFlush delivers higher throughput, significantly lower tail latency, and saves up to dozens of CPU cores per LSM-KVS server while reducing energy consumption.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get de92df78-96fe-4e6b-a2b5-e11f93180598Cited by top-tier papers2
- ArceKV: Towards Workload-driven LSM-compactions for Key-Value Store Under Dynamic WorkloadsJunfeng Liu, Haoxuan Xie, Siqiang LuoVLDB 2026
- Nezha: A Key-Value Separated Distributed Store with Optimized Raft IntegrationYangyang Wang, Yucong Dong, Ziqian Cheng, Zichen XuICDE 2026
Related papers
- DDS: DPU-optimized Disaggregated StorageQizhen Zhang, Philip A. Bernstein, Badrish Chandramouli, Jason Hu et al.VLDB 2024 · 12 citations
- FPGA-Accelerated Compactions for LSM-based Key-Value StoreTeng Zhang, Jianying Wang, Xuntao Cheng, Hao Xu et al.FAST 2020 · 99 citations
- Mitigating Resource Usage Dependency in Sorting-based KV Stores on Hybrid Storage Devices via Operation DecouplingQingyang Zhang, Yongkun Li, Yubiao Pan, Haoting Tang et al.USENIX ATC 2025 · 3 citations
- O3-LSM: Maximizing Disaggregated LSM Write Performance via Three-Layer OffloadingQi Lin, Gangqi Huang, Te Guo, Chang Guo et al.SIGMOD 2026 · 2 citations
- DShuffle: DPU-Optimized Shuffle Framework for Large-scale Data ProcessingChen Ding, Sicen Li, Kai Lu, Ting Yao et al.USENIX ATC 2025 · 2 citations
