Fast RDMA-based Ordered Key-Value Store using Remote Learned Cache
Xingda Wei, Rong Chen, Haibo Chen
摘要
RDMA ( Remote Direct Memory Access ) has gained considerable interests in network-attached in-memory key-value stores. However, traversing the remote tree-based index in ordered key-value stores with RDMA becomes a critical obstacle, causing an order-of-magnitude slowdown and limited scalability due to multiple round trips. Using index cache with conventional wisdom—caching partial data and traversing them locally—usually leads to limited effect because of unavoidable capacity misses, massive random accesses, and costly cache invalidations. We argue that the machine learning (ML) model is a perfect cache structure for the tree-based index, termed learned cache . Based on it, we design and implement XStore , an RDMA-based ordered key-value store with a new hybrid architecture that retains a tree-based index at the server to perform dynamic workloads (e.g., inserts) and leverages a learned cache at the client to perform static workloads (e.g., gets and scans). The key idea is to decouple ML model retraining from index updating by maintaining a layer of indirection from logical to actual positions of key-value pairs. It allows a stale learned cache to continue predicting a correct position for a lookup key. XStore ensures correctness using a validation mechanism with a fallback path and further uses speculative execution to minimize the cost of cache misses. Evaluations with YCSB benchmarks and production workloads show that a single XStore server can achieve over 80 million read-only requests per second. This number outperforms state-of-the-art RDMA-based ordered key-value stores (namely, DrTM-Tree, Cell, and eRPC+Masstree) by up to 5.9× (from 3.7×). For workloads with inserts, XStore still provides up to 3.5× (from 2.7×) throughput speedup, achieving 53M reqs/s. The learned cache can also reduce client-side memory usage and further provides an efficient memory-performance tradeoff, e.g., saving 99% memory at the cost of 20% peak throughput.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper39
- Sherman: A Write-Optimized Distributed B+Tree Index on Disaggregated MemoryQing Wang, Youyou Lu, Jiwu ShuSIGMOD 2022 · 被引用 99 次
- FINEdex: A Fine-grained Learned Index Scheme for Scalable and Concurrent Memory SystemsPengfei Li, Yu Hua, Jingnan Jia, Pengfei ZuoVLDB 2022 · 被引用 97 次
- ROLEX: A Scalable RDMA-oriented Learned Key-Value Store for Disaggregated Memory SystemsPengfei Li, Yu Hua, Pengfei Zuo, Zhangyu Chen 等FAST 2023 · 被引用 90 次
- Characterizing Off-path SmartNIC for Accelerating Distributed SystemsXingda Wei, Rongxin Cheng, Yuhan Yang, Rong Chen 等OSDI 2023 · 被引用 68 次
- GL-Cache: Group-level learning for efficient and high-performance cachingJuncheng Yang, Ziming Mao, Yao Yue, K. V. RashmiFAST 2023 · 被引用 60 次
它引用的顶会 Paper5
- ALEX: An Updatable Adaptive Learned IndexJialin Ding, Umar Farooq Minhas, Jia Yu, Chi Wang 等SIGMOD 2020 · 被引用 274 次
- Learning Multi-Dimensional IndexesVikram Nathan, Jialin Ding, Mohammad Alizadeh, Tim KraskaSIGMOD 2020 · 被引用 180 次
- From WiscKey to Bourbon: A Learned Index for Log-Structured Merge TreesYifan Dai, Yien Xu, Aishwarya Ganesan, Ramnatthan Alagappan 等OSDI 2020 · 被引用 138 次
- XIndex: a scalable learned index for multicore data storageChuzhe Tang, Youyun Wang, Zhiyuan Dong, Gansen Hu 等PPoPP 2020 · 被引用 109 次
- Cloudburst: Stateful Functions-as-a-ServiceVikram Sreekanti, Chenggang Wu, Xiayue Charles Lin, Johann Schleier-Smith 等VLDB 2020
相关 Paper
- RDMP-KV: designing remote direct memory persistence based key-value stores with PMEMTianxi Li, Dipti Shankar, Shashank Gugnani, Xiaoyi LuSC 2020 · 被引用 3 次
- ScaleStore: A Fast and Cost-Efficient Storage Engine using DRAM, NVMe, and RDMATobias Ziegler, Carsten Binnig, Viktor LeisSIGMOD 2022 · 被引用 51 次
- TeRM: Extending RDMA-Attached Memory with SSDZhe Yang, Qing Wang, Xiaojian Liao, Youyou Lu 等FAST 2024 · 被引用 6 次
- DPA-Store: An Ordered Network Data Path Key-Value StoreFrederic Schimmelpfennig, Jan Sass, Reza Salkhordeh, Martin Kröning 等OSDI 2026
- Cuckoo for Clients: Disaggregated Cuckoo HashingStewart Grant, Alex C. SnoerenUSENIX ATC 2025
