Limousine: Blending Learned and Classical Indexes to Self-Design Larger-than-Memory Cloud Storage Engines
Subarna Chatterjee, Mark F. Pekala, Lev Kruglyak, Stratos Idreos
Abstract
We present Limousine, a self-designing key-value storage engine, that can automatically morph to the nearoptimal storage engine architecture shape given a workload, a cloud budget, and a target performance. At its core, Limousine identifies the fundamental design principles of storage engines as combinations of learned and classical data structures that collaborate through algorithms for data storage and access. By unifying these principles over diverse hardware and three major cloud providers (AWS, GCP, and Azure), Limousine creates a massive design space of quindecillion (10 48 ) storage engine designs the vast majority of which do not exist in literature or industry. Limousine contains a distribution-aware IO model to accurately evaluate any candidate design. Using these models, Limousine searches within the exhaustive design space to construct a navigable continuum of designs connected along a Pareto frontier of cloud cost and performance. If storage engines contain learned components, Limousine also introduces efficient lazy write algorithms to optimize the holistic read-write performance. Once the near-optimal design is decided for the given context, Limousine automatically materializes the corresponding design in Rust code. Using the YCSB benchmark, we demonstrate that storage engines automatically designed and generated by Limousine scale better by up to 3 orders of magnitude when compared with state-of-the-art industry-leading engines such as RocksDB, WiredTiger, FASTER, and Cosine, over diverse workloads, data sets, and cloud budgets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Rethinking The Compaction Policies in LSM-treesHengrui Wang, Jiansheng Qiu, Fangzhou Yuan, Huanchen ZhangSIGMOD 2025 · 9 citations
- DobLIX: A Dual-Objective Learned Index for Log-Structured Merge TreesAlireza Heidari, Amirhossein Ahmadi, Wei ZhangVLDB 2025 · 4 citations
- HIRE: A Hybrid Learned Index for Robust and Efficient Performance under Mixed WorkloadsXinyi Zhang, Liang Liang, Anastasia Ailamaki, Jianliang XuSIGMOD 2026 · 2 citations
- ArceKV: Towards Workload-driven LSM-compactions for Key-Value Store Under Dynamic WorkloadsJunfeng Liu, Haoxuan Xie, Siqiang LuoVLDB 2026
Builds on10
- ALEX: An Updatable Adaptive Learned IndexJialin Ding, Umar Farooq Minhas, Jia Yu, Chi Wang et al.SIGMOD 2020 · 274 citations
- The PGM-index: a fully-dynamic compressed learned index with provable worst-case boundsPaolo Ferragina, Giorgio VinciguerraVLDB 2020 · 178 citations
- FINEdex: A Fine-grained Learned Index Scheme for Scalable and Concurrent Memory SystemsPengfei Li, Yu Hua, Jingnan Jia, Pengfei ZuoVLDB 2022 · 97 citations
- Learned Index: A Comprehensive Experimental EvaluationZhaoyan Sun, Xuanhe Zhou, Guoliang LiVLDB 2023 · 87 citations
- Constructing and Analyzing the LSM Compaction Design SpaceSubhadeep Sarkar, Dimitris Staratzis, Zichen Zhu, Manos AthanassoulisVLDB 2021 · 73 citations
Related papers
- Cosine: A Cloud-Cost Optimized Self-Designing Key-Value Storage EngineSubarna Chatterjee, Meena Jagadeesan, Wilson Qin, Stratos IdreosVLDB 2022 · 17 citations
- Dynamic read & write optimization with TurtleKVTony Astolfi, Vidya Silai, Darby Huye, Lan Liu et al.VLDB 2026
- Keigo: Co-designing Log-Structured Merge Key-Value Stores with a Non-Volatile, Concurrency-aware Storage HierarchyRúben Adão, Zhongjie Wu, Changjun Zhou, Oana Balmau et al.VLDB 2025
- Making LSM-Tree-based Key-Value Store Practical and Efficient for Multi-Tenant Serverless Cloud DatabasesYingjia Wang, Caixin Gong, Guoyun Zhu, Sheng Wang et al.SIGMOD 2026 · 1 citation
- Endure: A Robust Tuning Paradigm for LSM Trees Under Workload UncertaintyAndy Huynh, Harshal A. Chaudhari, Evimaria Terzi, Manos AthanassoulisVLDB 2022 · 28 citations
