AXE: A Task Decomposition Approach to Learned LSM Tuning
Andy Huynh, Anwesha Saha, Harshal A. Chaudhari, Manos Athanassoulis
Abstract
Log-Structured Merge (LSM) trees are used as the data structure of choice for key-value stores supporting a wide variety of applications. A common challenge for LSM-based systems is tuning them effectively, particularly as the complexity and number of tuning knobs increase. Prior work relies on expert-created cost models and expert-configured numerical solvers to produce high-quality tunings; however, these methods do not address tuning multiple instances at scale for various execution environments. On the other hand, using iterative learning, such as Bayesian Optimization (BO), relaxes the requirements for domain expertise and provides generalizability; however, it comes at a high cost, as it involves learning directly from database executions at deployment time. Furthermore, both approaches struggle with categorical tuning knobs that create a hard-to-navigate optimization space.
To address these challenges, we introduce AXE, a novel learned LSM tuning paradigm that decomposes the tuning task into two steps. First, AXE trains a learned cost model using existing performance modeling or execution logs, acting as a surrogate cost function in the tuning process. Second, AXE efficiently generates arbitrarily many training samples for a learned tuner optimized to identify high-performance tunings using the learned cost model as its loss function. This task decomposition approach generalizes well for tuning simple and complex LSM designs and requires no retraining, allowing AXE to be used for tuning at scale. Compared to BO, AXE recommends higher performing tunings than BO 71% of the time while incurring 100× smaller tuning overhead. We further show that AXE requires less domain knowledge to produce optimal tunings than traditional expert-configured tuning pipelines. Lastly, we compare AXE to both state-of-the-art machine learning methods and analytical methods to show that AXE outperforms all other LSM tuning baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c3f3cad9-03e2-4565-a01d-73688232f959Builds on12
- An End-to-End Learning-based Cost EstimatorJi Sun, Guoliang LiVLDB 2020 · 251 citations
- ResTune: Resource Oriented Tuning Boosted by Meta-Learning for Cloud DatabasesXinyi Zhang, Hong Wu, Zhuo Chang, Shuowei Jin et al.SIGMOD 2021 · 113 citations
- Cost Models for Big Data Query Processing: Learning, Retrofitting, and Our FindingsTarique Siddiqui, Alekh Jindal, Shi Qiao, Hiren Patel et al.SIGMOD 2020 · 80 citations
- Efficient Deep Learning Pipelines for Accurate Cost Estimations Over Large Scale Query WorkloadJohan Kok Zhi Kang, Gaurav, Sien Yi Tan, Feng Cheng et al.SIGMOD 2021 · 31 citations
- Endure: A Robust Tuning Paradigm for LSM Trees Under Workload UncertaintyAndy Huynh, Harshal A. Chaudhari, Evimaria Terzi, Manos AthanassoulisVLDB 2022 · 28 citations
Related papers
- CAMAL: Optimizing LSM-trees via Active LearningWeiping Yu, Siqiang Luo, Zihao Yu, Gao CongSIGMOD 2025 · 11 citations
- MCTuner: Spatial Decomposition-Enhanced Database Tuning via LLM-Guided ExplorationZihan Yan, Rui Xi, Mengshu HouSIGMOD 2026 · 3 citations
- DobLIX: A Dual-Objective Learned Index for Log-Structured Merge TreesAlireza Heidari, Amirhossein Ahmadi, Wei ZhangVLDB 2025 · 4 citations
- From WiscKey to Bourbon: A Learned Index for Log-Structured Merge TreesYifan Dai, Yien Xu, Aishwarya Ganesan, Ramnatthan Alagappan et al.OSDI 2020 · 138 citations
- LeaderKV: Improving Read Performance of KV Stores via Learned Index and Decoupled KV TableYi Wang, Jianan Yuan, Shangyu Wu, Huan Liu et al.ICDE 2024 · 12 citations
