RubbleDB: CPU-Efficient Replication with NVMe-oF
Haoyu Li, Sheng Jiang, Chen Chen, Ashwini Raina, Xingyu Zhu, Changxu Luo, Asaf Cidon
Abstract
Due to the need to perform expensive background compaction operations, the CPU is often a performance bottleneck of persistent key-value stores. In the case of replicated storage systems, which contain multiple identical copies of the data, we make the observation that CPU can be traded off for spare network bandwidth. Compactions can be executed only once, on one of the nodes, and the already-compacted data can be shipped to the other nodes' disks, saving them significant CPU time. In order to further drive down total CPU consumption, the file replication protocol can leverage NVMe-oF, a networked storage protocol that can offload the network and storage datapaths entirely to the NIC, requiring zero involvement from the target node's CPU. However, since NVMe-oF is a one-sided protocol, if used naively, it can easily cause data corruption or data loss at the target nodes.
We design RubbleDB, the first key-value store that takes advantage of NVMe-oF for efficient replication. RubbleDB introduces several novel design mechanisms that address the challenges of using NVMe-oF for replicated data, including pre-allocation of static files, a novel file metadata mapping mechanism, and a new method that enforces the order of applying version edits across replicas. These ideas can be applied to other settings beyond key-value stores, such as distributed file and backup systems. We implement RubbleDB on top of RocksDB and show it provides consistent CPU savings and increases throughput by up to 1.9× and reduces tail latency by up to 93.4% for write-heavy workloads, compared to replicated key-value stores, such as ZippyDB, which conduct compactions on all replica nodes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 93612522-394f-4c82-97cd-bb7293a07143Cited by top-tier papers6
- GeminiFS: A Companion File System for GPUsShi Qiu, Weinan Liu, Yifan Hu, Jianqin Yan et al.FAST 2025 · 17 citations
- Why Files If You Have a DBMS?Lam-Duy Nguyen, Viktor LeisICDE 2024 · 7 citations
- I/O in a Flash: Evolution of ONTAP to Low-Latency SSDsMatthew Curtis-Maury, Ram Kesavan, Bharadwaj V. R., Nikhil Mattankot et al.FAST 2024 · 3 citations
- Holistic and Automated Task Scheduling for Distributed LSM-tree-based StorageYuanming Ren, Siyuan Sheng, Zhang Cao, Yongkun Li et al.FAST 2026 · 1 citation
- CETOFS: A High-Performance File System with Host-Server Collaboration for Remote StorageWenqing Jia, Dejun Jiang, Jin XiongFAST 2026
Builds on8
- A large scale analysis of hundreds of in-memory cache clusters at TwitterJuncheng Yang, Yao Yue, K. V. RashmiOSDI 2020 · 245 citations
- Disaggregating Persistent Memory and Controlling Them Remotely: An Exploration of Passive Disaggregated Key-Value StoresShin-Yeh Tsai, Yizhou Shan, Yiying ZhangUSENIX ATC 2020 · 159 citations
- Building An Elastic Query Engine on Disaggregated StorageMidhul Vuppalapati, Justin Miron, Rachit Agarwal, Dan Truong et al.NSDI 2020 · 142 citations
- Differentiated Key-Value Storage Management for Balanced I/O PerformanceYongkun Li, Zhen Liu, Patrick P. C. Lee, Jiayu Wu et al.USENIX ATC 2021 · 79 citations
- Hailstorm: Disaggregated Compute and Storage for Distributed LSM-based DatabasesLaurent Bindschaedler, Ashvin Goel, Willy ZwaenepoelASPLOS 2020 · 51 citations
Related papers
- SpanDB: A Fast, Cost-Effective LSM-tree Based KV Store on Hybrid StorageHao Chen, Chaoyi Ruan, Cheng Li, Xiaosong Ma et al.FAST 2021 · 120 citations
- SplinterDB: Closing the Bandwidth Gap for NVMe Key-Value StoresAlexander Conway, Abhishek Gupta, Vijay Chidambaram, Martin Farach-Colton et al.USENIX ATC 2020 · 90 citations
- Light-Dedup: A Light-weight Inline Deduplication Framework for Non-Volatile Memory File SystemsJiansheng Qiu, Yanqi Pan, Wen Xia, Xiaojia Huang et al.USENIX ATC 2023 · 20 citations
- IONIA: High-Performance Replication for Modern Disk-based KV StoresYi Xu, Henry Zhu, Prashant Pandey, Alex Conway et al.FAST 2024 · 13 citations
- ListDB: Union of Write-Ahead Logs and Persistent SkipLists for Incremental Checkpointing on Persistent MemoryWonbae Kim, Chanyeol Park, Dongui Kim, Hyeongjun Park et al.OSDI 2022 · 47 citations
