AXE: A Task Decomposition Approach to Learned LSM Tuning
Andy Huynh, Anwesha Saha, Harshal A. Chaudhari, Manos Athanassoulis
摘要
Log-Structured Merge (LSM) trees are used as the data structure of choice for key-value stores supporting a wide variety of applications. A common challenge for LSM-based systems is tuning them effectively, particularly as the complexity and number of tuning knobs increase. Prior work relies on expert-created cost models and expert-configured numerical solvers to produce high-quality tunings; however, these methods do not address tuning multiple instances at scale for various execution environments. On the other hand, using iterative learning, such as Bayesian Optimization (BO), relaxes the requirements for domain expertise and provides generalizability; however, it comes at a high cost, as it involves learning directly from database executions at deployment time. Furthermore, both approaches struggle with categorical tuning knobs that create a hard-to-navigate optimization space.
To address these challenges, we introduce AXE, a novel learned LSM tuning paradigm that decomposes the tuning task into two steps. First, AXE trains a learned cost model using existing performance modeling or execution logs, acting as a surrogate cost function in the tuning process. Second, AXE efficiently generates arbitrarily many training samples for a learned tuner optimized to identify high-performance tunings using the learned cost model as its loss function. This task decomposition approach generalizes well for tuning simple and complex LSM designs and requires no retraining, allowing AXE to be used for tuning at scale. Compared to BO, AXE recommends higher performing tunings than BO 71% of the time while incurring 100× smaller tuning overhead. We further show that AXE requires less domain knowledge to produce optimal tunings than traditional expert-configured tuning pipelines. Lastly, we compare AXE to both state-of-the-art machine learning methods and analytical methods to show that AXE outperforms all other LSM tuning baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- An End-to-End Learning-based Cost EstimatorJi Sun, Guoliang LiVLDB 2020 · 被引用 251 次
- ResTune: Resource Oriented Tuning Boosted by Meta-Learning for Cloud DatabasesXinyi Zhang, Hong Wu, Zhuo Chang, Shuowei Jin 等SIGMOD 2021 · 被引用 113 次
- Cost Models for Big Data Query Processing: Learning, Retrofitting, and Our FindingsTarique Siddiqui, Alekh Jindal, Shi Qiao, Hiren Patel 等SIGMOD 2020 · 被引用 80 次
- Efficient Deep Learning Pipelines for Accurate Cost Estimations Over Large Scale Query WorkloadJohan Kok Zhi Kang, Gaurav, Sien Yi Tan, Feng Cheng 等SIGMOD 2021 · 被引用 31 次
- Endure: A Robust Tuning Paradigm for LSM Trees Under Workload UncertaintyAndy Huynh, Harshal A. Chaudhari, Evimaria Terzi, Manos AthanassoulisVLDB 2022 · 被引用 28 次
相关 Paper
- CAMAL: Optimizing LSM-trees via Active LearningWeiping Yu, Siqiang Luo, Zihao Yu, Gao CongSIGMOD 2025 · 被引用 11 次
- MCTuner: Spatial Decomposition-Enhanced Database Tuning via LLM-Guided ExplorationZihan Yan, Rui Xi, Mengshu HouSIGMOD 2026 · 被引用 3 次
- DobLIX: A Dual-Objective Learned Index for Log-Structured Merge TreesAlireza Heidari, Amirhossein Ahmadi, Wei ZhangVLDB 2025 · 被引用 4 次
- From WiscKey to Bourbon: A Learned Index for Log-Structured Merge TreesYifan Dai, Yien Xu, Aishwarya Ganesan, Ramnatthan Alagappan 等OSDI 2020 · 被引用 138 次
- LeaderKV: Improving Read Performance of KV Stores via Learned Index and Decoupled KV TableYi Wang, Jianan Yuan, Shangyu Wu, Huan Liu 等ICDE 2024 · 被引用 12 次
