Pome: Parallelizing I/Os and Computations for Efficient LSM-tree-based Data Storage
Yanpeng Hu, Li Zhu, Lei Jia, Chundong Wang
Abstract
CPU computations and I/O operations are fundamental to data storage systems. Storage systems conduct computations with their user threads, such as sorting data for orderliness. They handle I/Os through system calls (syscalls) including file write, read, and fsync, which the OS's kernel threads perform with storage devices.
Today, LSM-tree-based storage systems are widely deployed in production environments. Compaction is an essential operation that LSM-tree employs to maintain its tiered tree-like structure by re-sorting and re-storing data through computations and I/Os, respectively. In this paper, we first overhaul the procedure of a compaction. We find that computations and I/Os execute in a sequential order. After re-sorting data, the user thread waits for a kernel thread to complete file write and fsync I/Os. These costly synchronous I/Os create a severely long critical path that affects the performance of LSM-tree. To address this issue, we propose parallelizing I/Os and computations for efficient LSM-tree-based data storage (Pome). Pome decouples computations from I/Os within each compaction by referring to its new protocol that moves I/O operations out of the critical path. To this end, it conducts asynchronous I/Os by using io_uring. Furthermore, regarding the potential I/O congestion caused by accelerated compactions, Pome incorporates an adaptive I/O rate limiter to achieve smooth execution. We prototype Pome on top of RocksDB. Experimental results demonstrate that Pome significantly improves the performance of RocksDB and outperforms several state-of-the-art LSM-tree variants.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4481319f-ccd5-4525-9b77-8e16aa0fbe1aCited by top-tier papers1
Ask how each one uses itBuilds on12
- SpanDB: A Fast, Cost-Effective LSM-tree Based KV Store on Hybrid StorageHao Chen, Chaoyi Ruan, Cheng Li, Xiaosong Ma et al.FAST 2021 · 120 citations
- XRP: In-Kernel Storage Functions with eBPFYuhong Zhong, Haoyu Li, Yu Jian Wu, Ioannis Zarkadas et al.OSDI 2022 · 100 citations
- FPGA-Accelerated Compactions for LSM-based Key-Value StoreTeng Zhang, Jianying Wang, Xuntao Cheng, Hao Xu et al.FAST 2020 · 99 citations
- AC-Key: Adaptive Caching for LSM-based Key-Value StoresFenggang Wu, Ming-Hong Yang, Baoquan Zhang, David H. C. DuUSENIX ATC 2020 · 81 citations
- ADOC: Automatically Harmonizing Dataflow Between Components in Log-Structured Key-Value Stores for Improved PerformanceJinghuan Yu, Sam H. Noh, Young-ri Choi, Chun Jason XueFAST 2023 · 52 citations
Related papers
- Resystance: Unleashing Hidden Performance of Compaction in LSM-Trees Via eBPFHongsu Byun, Seungjae Lee, Honghyeon Yoo, Myoungjoon Kim et al.ICDE 2026
- Reducing Write Amplification of LSM-Tree with Block-Grained CompactionXiaoliang Wang, Peiquan Jin, Bei Hua, Hai Long et al.ICDE 2022 · 26 citations
- Rethinking The Compaction Policies in LSM-treesHengrui Wang, Jiansheng Qiu, Fangzhou Yuan, Huanchen ZhangSIGMOD 2025 · 9 citations
- Closing the Performance Gap between Leveling and Tiering Compaction via Bundle CompactionRuicheng Liu, Peiquan Jin, Xiaoliang Wang, Yongping Luo et al.HPDC 2023 · 5 citations
- Constructing and Analyzing the LSM Compaction Design SpaceSubhadeep Sarkar, Dimitris Staratzis, Zichen Zhu, Manos AthanassoulisVLDB 2021 · 73 citations
