Tetris: Lightweight Hyperparameter Auto-Tuning for Mitigating Performance Spikes in LSM-KVS
Yina Lv, Wenhao Zhu, Qiao Li, Quanqing Xu, Congming Gao, Chuanhui Yang, Xiaoli Wang, Chun Jason Xue
摘要
LSM-trees have been widely adopted in modern database systems owing to their log-structured design and sequential write efficiency. This design makes them particularly suitable for write-intensive, large-scale scenarios. However, a critical challenge lies in the flush and compaction processes in the LSM-tree, which often lead to performance fluctuation. Existing tuning strategies, including manual configurations and machine learning or LLM-based methods, struggle to adapt to dynamic workload patterns (e.g., sequential vs. random writes, Zipfian vs. uniform distributions). These methods also incur high tuning overhead and fail to mitigate severe I/O contention and performance spikes. In this paper, we propose Tetris, a lightweight hyperparameter auto-tuning framework designed to mitigate performance spikes in LSM-based key-value stores. Tetris dynamically adjusts LSM-tree configurable parameters by monitoring performance spikes, workload patterns, and realtime resource utilization during runtime. It contains three key components: a performance-driven tuning trigger, which determines when to initiate parameter adjustments; a workload-aware parameter selector, which identifies what parameters to tune based on workload characteristics; and a resource-efficient autotuner, which optimizes resource utilization while adapting to the current LSM-tree status. We evaluate Tetris on RocksDB using db_bench benchmarks and real-world YCSB workloads. Experimental results show that Tetris mitigates performance spikes across various workloads with minimal overhead. Compared to ADOC, it reduces average latency by 18% (up to 62%) and improves throughput by 38% (up to 1.64×). For tail latency, Tetris cuts P99 and P99.9 read latency by 54% and 53%, respectively, while maintaining comparable write tail latency with a 66% lower standard deviation in write latency.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- ArceKV: Towards Workload-driven LSM-compactions for Key-Value Store Under Dynamic WorkloadsJunfeng Liu, Haoxuan Xie, Siqiang LuoVLDB 2026
- ADOC: Automatically Harmonizing Dataflow Between Components in Log-Structured Key-Value Stores for Improved PerformanceJinghuan Yu, Sam H. Noh, Young-ri Choi, Chun Jason XueFAST 2023 · 被引用 52 次
- Dynamic read & write optimization with TurtleKVTony Astolfi, Vidya Silai, Darby Huye, Lan Liu 等VLDB 2026
- Rethinking The Compaction Policies in LSM-treesHengrui Wang, Jiansheng Qiu, Fangzhou Yuan, Huanchen ZhangSIGMOD 2025 · 被引用 9 次
- Holistic and Automated Task Scheduling for Distributed LSM-tree-based StorageYuanming Ren, Siyuan Sheng, Zhang Cao, Yongkun Li 等FAST 2026 · 被引用 1 次
