Rethinking Learned Cost Models: Why Start from Scratch?
Jiani Yang, Sai Wu, Dongxiang Zhang, Jian Dai, Feifei Li, Gang Chen
Abstract
Recent work has applied learning-based approaches to replace the conventional cost model, but these approaches are expensive to train and result in high inference overheads. Furthermore, due to a lack of explainability, models trained for one database may not be easily transferred to another, requiring a complete re-training process. In this paper, we propose a new approach to tuning the conventional formula-based cost model for DBMS. Our approach involves identifying important parameters within the cost model rules and using a fast-learning model to adjust them for each specific hardware and software configuration of the DBMS deployment. We dynamically partition the search space of hardware and software configurations to gradually refine the cost model estimation. To apply our cost model to a new DBMS instance, we start with a rough estimation and progressively refine it with finer granularity. Our experiments with different hardware and software configurations show that our approach enables the conventional cost model to be quickly transferred to any database instance, achieving comparable results to a fine-tuned learning-based model. Overall, our approach provides a practical solution to tuning the conventional cost model for DBMS, with significant benefits in terms of reduced cost and improved performance.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers6
- λ-Tune: Harnessing Large Language Models for Automated Database System TuningVictor Giannakouris, Immanuel TrummerSIGMOD 2025 · 20 citations
- How Good are Learned Cost Models, Really? Insights from Query Optimization TasksRoman Heinrich, Manisha Luthra, Johannes Wehrstein, Harald Kornmayer et al.SIGMOD 2025 · 13 citations
- T3: Accurate and Fast Performance Prediction for Relational Database Systems With Compiled Decision TreesMaximilian Rieger, Thomas NeumannSIGMOD 2025 · 5 citations
- Breaking the Isolation-Freshness Trade-off: Joint Adaptive Storage Optimization for HTAP SystemsZhenghao Ding, Xinyi Zhang, Chao Zhang, Yishen Sun et al.VLDB 2026 · 1 citation
- OBELISK: Efficient Offline Query Planning with Bayesian Optimization-Informed Language Model ReasoningZhicheng Pan, Wenwen Sun, Yuanjia Zhang, Terence Purcell et al.VLDB 2026
Related papers
- DISTILL: Low-Overhead Data-Driven Techniques for Filtering and Costing Indexes for Scalable Index TuningTarique Siddiqui, Wentao Wu, Vivek R. Narasayya, Surajit ChaudhuriVLDB 2022 · 36 citations
- Zero-Shot Cost Models for Out-of-the-box Learned Cost PredictionBenjamin Hilprecht, Carsten BinnigVLDB 2022 · 90 citations
- Explainable Database Management System Configuration Tuning through CounterfactualsXinyue Shao, Hongzhi Wang, Xiao Zhu, Tianyu Mu et al.ICDE 2024
- CGPTuner: a Contextual Gaussian Process Bandit Approach for the Automatic Tuning of IT Configurations Under Varying Workload ConditionsStefano Cereda, Stefano Valladares, Paolo Cremonesi, Stefano DoniVLDB 2021 · 74 citations
- Crop: An Analytical Cost Model for Cross-Platform Performance Prediction of Tensor ProgramsXinyu Sun, Yu Zhang, Shuo Liu, Yi ZhaiDAC 2024 · 2 citations
