Characterizing, Modeling, and Benchmarking RocksDB Key-Value Workloads at Facebook
Zhichao Cao, Siying Dong, Sagar Vemuri, David H. C. Du
Abstract
Persistent key-value stores are widely used as building blocks in today's IT infrastructure for managing and storing large amounts of data. However, studies of characterizing real-world workloads for key-value stores are limited due to the lack of tracing/analyzing tools and the difficulty of collecting traces in operational environments. In this paper, we first present a detailed characterization of workloads from three typical RocksDB production use cases at Facebook: UDB (a MySQL storage layer for social graph data), ZippyDB (a distributed key-value store), and UP2X (a distributed key-value store for AI/ML services). These characterizations reveal several interesting findings: first, that the distribution of key and value sizes are highly related to the use cases/applications; second, that the accesses to key-value pairs have a good locality and follow certain special patterns; and third, that the collected performance metrics show a strong diurnal pattern in the UDB, but not the other two.
We further discover that although the widely used key-value benchmark YCSB provides various workload configurations and key-value pair access distribution models, the YCSBtriggered workloads for underlying storage systems are still not close enough to the workloads we collected due to ignorance of key-space localities. To address this issue, we propose a key-range based modeling and develop a benchmark that can better emulate the workloads of real-world key-value stores. This benchmark can synthetically generate more precise key-value queries that represent the reads and writes of key-value stores to the underlying storage system.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e5e4c1a0-5231-4ea4-bb08-23fca98cc305Cited by top-tier papers118
- ALEX: An Updatable Adaptive Learned IndexJialin Ding, Umar Farooq Minhas, Jia Yu, Chi Wang et al.SIGMOD 2020 · 274 citations
- A large scale analysis of hundreds of in-memory cache clusters at TwitterJuncheng Yang, Yao Yue, K. V. RashmiOSDI 2020 · 245 citations
- Tsunami: A Learned Multi-dimensional Index for Correlated Data and Skewed WorkloadsJialin Ding, Vikram Nathan, Mohammad Alizadeh, Tim KraskaVLDB 2021 · 178 citations
- SpanDB: A Fast, Cost-Effective LSM-tree Based KV Store on Hybrid StorageHao Chen, Chaoyi Ruan, Cheng Li, Xiaosong Ma et al.FAST 2021 · 120 citations
- Evolution of Development Priorities in Key-value Stores Serving Large-scale Applications: The RocksDB ExperienceSiying Dong, Andrew Kryczka, Yanqin Jin, Michael StummFAST 2021 · 110 citations
Related papers
- EvenDB: optimizing key-value storage for spatial localityEran Gilad, Edward Bortnikov, Anastasia Braginsky, Yonatan Gottesman et al.EuroSys 2020 · 29 citations
- A new benchmark harness for systematic and robust evaluation of streaming state storesEsmail Asyabi, Yuanli Wang, John Liagouris, Vasiliki Kalavri et al.EuroSys 2022 · 15 citations
- H-Rocks: CPU-GPU accelerated Heterogeneous RocksDB on Persistent MemoryShweta Pandey, Arkaprava BasuSIGMOD 2025 · 4 citations
- TAOBench: An End-to-End Benchmark for Social Networking WorkloadsAudrey Cheng, Xiao Shi, Aaron N. Kabcenell, Shilpa Lawande et al.VLDB 2022 · 20 citations
- AnyKey: A Key-Value SSD for All Workload TypesChanyoung Park, Jungho Lee, Chun-Yi Liu, Kyungtae Kang et al.ASPLOS 2025 · 2 citations
