How Much Can RocksDB Chew? Achieving Near-Zero Write Stalls with Sustainable RocksDB
Hojin Shin, Yongmin Lee, Seehwan Yoo, Jongmoo Choi
Abstract
Modern data-intensive applications, from microservices to realtime AI serving, demand consistently low tail latency from backend storage. However, log-structured merge-tree-based key-value stores like RocksDB are structurally prone to unpredictable write stalls. These stalls stem from a fundamental architectural decoupling of foreground write ingress and background data reorganization. By design, the system absorbs foreground writes at maximum speed without monitoring its actual time-varying compaction capacity. As a result, it accumulates internal pressure until rigid capacity thresholds are breached, triggering reactive safeguards that abruptly freeze all foreground writes. Relying on this reactive "stop-and-go" approach induces a persistent limit-cycle behavior, undermining long-run predictability and strict latency guarantees.
We reframe write stalls as a continuous control problem. S-RocksDB is a sustainable admission controller that regulates foreground ingress to match the system's time-varying compaction capacity. Since this capacity varies at runtime, S-RocksDB employs online reinforcement learning to discover a sustainable admission rate. To ensure safe learning, a three-state operational model (SAFE, SEMI-SAFE, UNSAFE) confines exploration to stable conditions and deploys deterministic guardrails before stalls can occur. In 24-hour evaluations, S-RocksDB reduces over 64.3M stalled writes to just 69, bounds P99.99 tail latency to sub-0.11 ms, and delivers predictable throughput with the lowest resource footprint among all compared systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ac6e0d4d-3038-4157-ba1c-bd24cfb1e597Builds on14
- ALEX: An Updatable Adaptive Learned IndexJialin Ding, Umar Farooq Minhas, Jia Yu, Chi Wang et al.SIGMOD 2020 · 274 citations
- ZNS: Avoiding the Block Interface Tax for Flash-based SSDsMatias Bjørling, Abutalib Aghayev, Hans Holmberg, Aravind Ramesh et al.USENIX ATC 2021 · 221 citations
- From WiscKey to Bourbon: A Learned Index for Log-Structured Merge TreesYifan Dai, Yien Xu, Aishwarya Ganesan, Ramnatthan Alagappan et al.OSDI 2020 · 138 citations
- LinnOS: Predictability on Unpredictable Flash Storage with a Light Neural NetworkMingzhe Hao, Levent Toksoz, Nanqinqin Li, Edward Edberg Halim et al.OSDI 2020 · 97 citations
- SplinterDB: Closing the Bandwidth Gap for NVMe Key-Value StoresAlexander Conway, Abhishek Gupta, Vijay Chidambaram, Martin Farach-Colton et al.USENIX ATC 2020 · 90 citations
Related papers
- ADOC: Automatically Harmonizing Dataflow Between Components in Log-Structured Key-Value Stores for Improved PerformanceJinghuan Yu, Sam H. Noh, Young-ri Choi, Chun Jason XueFAST 2023 · 52 citations
- Learning to Optimize LSM-trees: Towards A Reinforcement Learning based Key-Value Store for Dynamic WorkloadsDingheng Mo, Fanchao Chen, Siqiang Luo, Caihua ShanSIGMOD 2024 · 26 citations
- ArceKV: Towards Workload-driven LSM-compactions for Key-Value Store Under Dynamic WorkloadsJunfeng Liu, Haoxuan Xie, Siqiang LuoVLDB 2026
- Reinforcement Learning-Assisted Cache Cleaning to Mitigate Long-Tail Latency in DM-SMRYungang Pan, Zhiping Jia, Zhaoyan Shen, Bingzhe Li et al.DAC 2021 · 12 citations
- Tidehunter: Large-Value Storage With Minimal Data RelocationAndrey Chursin, Lefteris Kokoris-Kogias, Alex Orlov, Alberto Sonnino et al.VLDB 2026 · 1 citation
