A new benchmark harness for systematic and robust evaluation of streaming state stores
Esmail Asyabi, Yuanli Wang, John Liagouris, Vasiliki Kalavri, Azer Bestavros
摘要
Modern stream processing systems often rely on embedded key-value stores, like RocksDB, to manage the state of longrunning computations. Evaluating the performance of these stores when used for streaming workloads is cumbersome as it requires the configuration and deployment of a stream processing system that integrates the respective store, and the execution of representative queries to collect measurements.
To address this issue, in this paper, we start with an empirical characterization of streaming state access workloads collected from Apache Flink and RocksDB, using three publicly available datasets, and we show that the characteristics of real traces cannot be approximated with existing benchmarks. Next, we present Gadget, a new benchmark harness that generates realistic streaming state access workloads to enable easy and thorough performance evaluation of standalone KV stores through accurate simulation of streaming operator logic. Finally, we use Gadget to investigate the suitability of RocksDB as the de facto kv store for stream processing systems. Interestingly, we find that, although RocksDB provides robust results, it is outperformed by FASTER and BerkeleyDB in six out of eleven workloads. Our results reveal a wide performance gap between the current performance of streaming state stores and what could be achieved with workload-aware approaches.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Accelerating String-key Learned Index Structures via Memoization-based Incremental TrainingMinsu Kim, Jinwoo Hwang, Guseul Heo, Seiyeon Cho 等VLDB 2024 · 被引用 10 次
- CAPSys: Contention-aware task placement for data stream processingYuanli Wang, Lei Huang, Zikun Wang, Vasiliki Kalavri 等EuroSys 2025 · 被引用 8 次
- Low-Latency Stateful Stream Processing Through Timely and Accurate PrefetchingEleni Zapridou, Anastasia AilamakiICDE 2026
它引用的顶会 Paper4
- A large scale analysis of hundreds of in-memory cache clusters at TwitterJuncheng Yang, Yao Yue, K. V. RashmiOSDI 2020 · 被引用 245 次
- The CacheLib Caching Engine: Design and Experiences at ScaleBenjamin Berg, Daniel S. Berger, Sara McAllister, Isaac Grosof 等OSDI 2020 · 被引用 145 次
- Lethe: A Tunable Delete-Aware LSM EngineSubhadeep Sarkar, Tarikul Islam Papon, Dimitris Staratzis, Manos AthanassoulisSIGMOD 2020 · 被引用 68 次
- Characterizing, Modeling, and Benchmarking RocksDB Key-Value Workloads at FacebookZhichao Cao, Siying Dong, Sagar Vemuri, David H. C. DuFAST 2020
相关 Paper
- FlowKV: A Semantic-Aware Store for Large-Scale State Management of Stream Processing EnginesGyewon Lee, Jaewoo Maeng, Jinsol Park, Jangho Seo 等EuroSys 2023 · 被引用 6 次
- H-Rocks: CPU-GPU accelerated Heterogeneous RocksDB on Persistent MemoryShweta Pandey, Arkaprava BasuSIGMOD 2025 · 被引用 4 次
- How Reliable Are Streams? End-to-End Processing-Guarantee Validation and Performance Benchmarking of Stream Processing SystemsJawad Tahir, Ruben Mayer, Christoph Doblander, Hans-Arno JacobsenVLDB 2025 · 被引用 3 次
- ArceKV: Towards Workload-driven LSM-compactions for Key-Value Store Under Dynamic WorkloadsJunfeng Liu, Haoxuan Xie, Siqiang LuoVLDB 2026
- gParaKV: A GPGPU-accelerated Key-Value Separation-based KV Store with Optimized Compaction and Garbage CollectionHui Sun, Xiangxiang Jiang, Xiao Qin, Song Jiang 等SC 2025 · 被引用 3 次
