Evolution of Development Priorities in Key-value Stores Serving Large-scale Applications: The RocksDB Experience
Siying Dong, Andrew Kryczka, Yanqin Jin, Michael Stumm
Abstract
RocksDB is a key-value store targeting large-scale distributed systems and optimized for Solid State Drives (SSDs). This paper describes how our priorities in developing RocksDB have evolved over the last eight years. The evolution is the result both of hardware trends and of extensive experience running RocksDB at scale in production at a number of organizations. We describe how and why RocksDB's resource optimization target migrated from write amplification, to space amplification, to CPU utilization. Lessons from running large-scale applications taught us that resource allocation needs to be managed across different RocksDB instances, that data format needs to remain backward and forward compatible to allow incremental software rollout, and that appropriate support for database replication and backups are needed. Lessons from failure handling taught us that data corruption errors needed to be detected earlier and at every layer of the system.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2f5aca87-f9eb-425d-b5dc-06279aa0b1deCited by top-tier papers26
- Spooky: Granulating LSM-Tree Compactions CorrectlyNiv Dayan, Tamar Weiss, Shmuel Dashevsky, Michael Pan et al.VLDB 2022 · 57 citations
- Understanding Silent Data Corruptions in a Large Production CPU PopulationShaobu Wang, Guangyan Zhang, Junyu Wei, Yang Wang et al.SOSP 2023 · 56 citations
- ADOC: Automatically Harmonizing Dataflow Between Components in Log-Structured Key-Value Stores for Improved PerformanceJinghuan Yu, Sam H. Noh, Young-ri Choi, Chun Jason XueFAST 2023 · 52 citations
- Pacman: An Efficient Compaction Approach for Log-Structured Key-Value Store on Persistent MemoryJing Wang, Youyou Lu, Qing Wang, Minhui Xie et al.USENIX ATC 2022 · 44 citations
- Near-Data Processing in Database Systems on Native Computational Storage under HTAP WorkloadsTobias Vinçon, Christian Knödler, Leonardo Solis-Vasquez, Arthur Bernhardt et al.VLDB 2022 · 34 citations
Builds on1
Related papers
- EvenDB: optimizing key-value storage for spatial localityEran Gilad, Edward Bortnikov, Anastasia Braginsky, Yonatan Gottesman et al.EuroSys 2020 · 29 citations
- H-Rocks: CPU-GPU accelerated Heterogeneous RocksDB on Persistent MemoryShweta Pandey, Arkaprava BasuSIGMOD 2025 · 4 citations
- ZoneKV: A Space-Efficient Key-Value Store for ZNS SSDsMingchen Lu, Peiquan Jin, Xiaoliang Wang, Yongping Luo et al.DAC 2023 · 16 citations
- SpanDB: A Fast, Cost-Effective LSM-tree Based KV Store on Hybrid StorageHao Chen, Chaoyi Ruan, Cheng Li, Xiaosong Ma et al.FAST 2021 · 120 citations
- SplinterDB: Closing the Bandwidth Gap for NVMe Key-Value StoresAlexander Conway, Abhishek Gupta, Vijay Chidambaram, Martin Farach-Colton et al.USENIX ATC 2020 · 90 citations
