Characterizing, Modeling, and Benchmarking RocksDB Key-Value Workloads at Facebook
Zhichao Cao, Siying Dong, Sagar Vemuri, David H. C. Du
摘要
Persistent key-value stores are widely used as building blocks in today's IT infrastructure for managing and storing large amounts of data. However, studies of characterizing real-world workloads for key-value stores are limited due to the lack of tracing/analyzing tools and the difficulty of collecting traces in operational environments. In this paper, we first present a detailed characterization of workloads from three typical RocksDB production use cases at Facebook: UDB (a MySQL storage layer for social graph data), ZippyDB (a distributed key-value store), and UP2X (a distributed key-value store for AI/ML services). These characterizations reveal several interesting findings: first, that the distribution of key and value sizes are highly related to the use cases/applications; second, that the accesses to key-value pairs have a good locality and follow certain special patterns; and third, that the collected performance metrics show a strong diurnal pattern in the UDB, but not the other two.
We further discover that although the widely used key-value benchmark YCSB provides various workload configurations and key-value pair access distribution models, the YCSBtriggered workloads for underlying storage systems are still not close enough to the workloads we collected due to ignorance of key-space localities. To address this issue, we propose a key-range based modeling and develop a benchmark that can better emulate the workloads of real-world key-value stores. This benchmark can synthetically generate more precise key-value queries that represent the reads and writes of key-value stores to the underlying storage system.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper118
- ALEX: An Updatable Adaptive Learned IndexJialin Ding, Umar Farooq Minhas, Jia Yu, Chi Wang 等SIGMOD 2020 · 被引用 274 次
- A large scale analysis of hundreds of in-memory cache clusters at TwitterJuncheng Yang, Yao Yue, K. V. RashmiOSDI 2020 · 被引用 245 次
- Tsunami: A Learned Multi-dimensional Index for Correlated Data and Skewed WorkloadsJialin Ding, Vikram Nathan, Mohammad Alizadeh, Tim KraskaVLDB 2021 · 被引用 178 次
- SpanDB: A Fast, Cost-Effective LSM-tree Based KV Store on Hybrid StorageHao Chen, Chaoyi Ruan, Cheng Li, Xiaosong Ma 等FAST 2021 · 被引用 120 次
- Evolution of Development Priorities in Key-value Stores Serving Large-scale Applications: The RocksDB ExperienceSiying Dong, Andrew Kryczka, Yanqin Jin, Michael StummFAST 2021 · 被引用 110 次
相关 Paper
- EvenDB: optimizing key-value storage for spatial localityEran Gilad, Edward Bortnikov, Anastasia Braginsky, Yonatan Gottesman 等EuroSys 2020 · 被引用 29 次
- A new benchmark harness for systematic and robust evaluation of streaming state storesEsmail Asyabi, Yuanli Wang, John Liagouris, Vasiliki Kalavri 等EuroSys 2022 · 被引用 15 次
- H-Rocks: CPU-GPU accelerated Heterogeneous RocksDB on Persistent MemoryShweta Pandey, Arkaprava BasuSIGMOD 2025 · 被引用 4 次
- TAOBench: An End-to-End Benchmark for Social Networking WorkloadsAudrey Cheng, Xiao Shi, Aaron N. Kabcenell, Shilpa Lawande 等VLDB 2022 · 被引用 20 次
- AnyKey: A Key-Value SSD for All Workload TypesChanyoung Park, Jungho Lee, Chun-Yi Liu, Kyungtae Kang 等ASPLOS 2025 · 被引用 2 次
