MFTune: An Efficient Multi-fidelity Framework for Spark SQL Configuration Tuning
Beicheng Xu, Lingching Tung, Yuchen Wang, Yupeng Lu, Bin Cui
Abstract
Apache Spark SQL is a cornerstone of modern big data analytics. However, optimizing Spark SQL performance is challenging due to its vast configuration space and the prohibitive cost of evaluating massive workloads. Existing tuning methods predominantly rely on full-fidelity evaluations, which are extremely time-consuming, often leading to suboptimal performance within practical budgets. While multi-fidelity optimization offers a potential solution, directly applying standard techniques-such as data volume reduction or early stopping-proves ineffective for Spark SQL as they fail to preserve performance correlations or represent true system bottlenecks. To address these challenges, we propose MFTune, an efficient multi-fidelity framework that introduces a query-based fidelity partitioning strategy, utilizing representative SQL subsets to provide accurate, low-cost proxies. To navigate the huge search space, MFTune incorporates a density-based optimization mechanism for automated knob and range compression, alongside an adapted transfer learning approach and a two-phase warm start to further accelerate the tuning process. Experimental results on TPC-H and TPC-DS benchmarks demonstrate that MFTune significantly outperforms five state-of-the-art tuning methods, identifying superior configurations within practical time constraints.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on23
- ResTune: Resource Oriented Tuning Boosted by Meta-Learning for Cloud DatabasesXinyi Zhang, Hong Wu, Zhuo Chang, Shuowei Jin et al.SIGMOD 2021 · 113 citations
- Lero: A Learning-to-Rank Query OptimizerRong Zhu, Wei Chen, Bolin Ding, Xingguang Chen et al.VLDB 2023 · 102 citations
- Facilitating Database Tuning with Hyper-Parameter Optimization: A Comprehensive Experimental EvaluationXinyi Zhang, Zhuo Chang, Yang Li, Hong Wu et al.VLDB 2022 · 88 citations
- LlamaTune: Sample-Efficient DBMS Configuration TuningKonstantinos Kanellis, Cong Ding, Brian Kroth, Andreas Müller et al.VLDB 2022 · 73 citations
- DSB: A Decision Support Benchmark for Workload-Driven and Traditional Database SystemsBailu Ding, Surajit Chaudhuri, Johannes Gehrke, Vivek R. NarasayyaVLDB 2021 · 62 citations
Related papers
- LOCAT: Low-Overhead Online Configuration Auto-Tuning of Spark SQL ApplicationsJinhan Xin, Kai Hwang, Zhibin YuSIGMOD 2022 · 34 citations
- Adaptive Code Learning for Spark Configuration TuningChen Lin, Junqing Zhuang, Jiadong Feng, Hui Li et al.ICDE 2022 · 28 citations
- A Spark Optimizer for Adaptive, Fine-Grained Parameter TuningChenghao Lyu, Qi Fan, Philippe Guyard, Yanlei DiaoVLDB 2024 · 9 citations
- MCTuner: Spatial Decomposition-Enhanced Database Tuning via LLM-Guided ExplorationZihan Yan, Rui Xi, Mengshu HouSIGMOD 2026 · 3 citations
- LEAP: A Low-cost Spark SQL Query Optimizer using Pairwise ComparisonJunhao Ye, Jiahui Li, Lu Chen, Yuren Mao et al.VLDB 2025 · 2 citations
