LOCAT: Low-Overhead Online Configuration Auto-Tuning of Spark SQL Applications
Jinhan Xin, Kai Hwang, Zhibin Yu
Abstract
Spark SQL has been widely deployed in industry but it is challenging to tune its performance. Recent studies try to employ machine learning (ML) to solve this problem. They however suffer from two drawbacks. First, it takes a long time (high overhead) to collect training samples. Second, the optimal configuration for one input data size of the same application might not be optimal for others.
To address these issues, we propose a novel Bayesian Optimization (BO) based approach named LOCAT to automatically tune the configurations of Spark SQL applications online. LOCAT innovates three techniques. The first technique, named QCSA, eliminates the configuration-insensitive queries by Query Configuration Sensitivity Analysis (QCSA) when collecting training samples. The second technique, dubbed DAGP, is a Datasize-Aware Gaussian Process (DAGP) which models the performance of an application as a distribution of functions of configuration parameters as well as input data size. The third technique, called IICP, Identifies Important Configuration Parameters (IICP) with respect to performance and only tunes the important parameters. As such, LOCAT can tune the configurations of a Spark SQL application with low overhead and adapt to different input data sizes.
We employ Spark SQL applications from benchmark suites๐ ๐๐ถ-๐ท๐, ๐ ๐๐ถ -๐ป , and ๐ป๐๐ต๐๐๐โ running on two significantly different clusters, a four-node ARM cluster and an eight-node x86 cluster, to evaluate LOCAT. The experimental results on the ARM cluster show that LOCAT accelerates the optimization procedures of Tuneful [22], DAC [66], , and QTune [37] by factors of 6.4ร, 7.0ร, 4.1ร, and 9.7ร on average, respectively. On the x86 cluster, LOCAT reduces the optimization time of Tuneful, DAC, GBO-RL, and QTune by factors of 6.4ร, 6.3ร, 4.0ร, and 9.2ร on average, respectively. Moreover, LOCAT improves the performance of the applications on
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 564d59b4-a3f6-458c-9171-5c65cbf62484Cited by top-tier papers10
- OPPerTune: Post-Deployment Configuration Tuning of Services Made EasyGagan Somashekar, Karan Tandon, Anush Kini, Chieh-Chun Chang et al.NSDI 2024 ยท 24 citations
- VDTuner: Automated Performance Tuning for Vector Data Management SystemsTiannuo Yang, Wen Hu, Wangqi Peng, Yusen Li et al.ICDE 2024 ยท 12 citations
- A Spark Optimizer for Adaptive, Fine-Grained Parameter TuningChenghao Lyu, Qi Fan, Philippe Guyard, Yanlei DiaoVLDB 2024 ยท 9 citations
- LEAP: A Low-cost Spark SQL Query Optimizer using Pairwise ComparisonJunhao Ye, Jiahui Li, Lu Chen, Yuren Mao et al.VLDB 2025 ยท 2 citations
- InferLog: Accelerating LLM Inference for Online Log Parsing via ICL-oriented Prefix CachingYilun Wang, Pengfei Chen, Haiyu Huang, Zilong He et al.ICSE 2026 ยท 1 citation
Builds on1
Related papers
- Adaptive Code Learning for Spark Configuration TuningChen Lin, Junqing Zhuang, Jiadong Feng, Hui Li et al.ICDE 2022 ยท 28 citations
- MFTune: An Efficient Multi-fidelity Framework for Spark SQL Configuration TuningBeicheng Xu, Lingching Tung, Yuchen Wang, Yupeng Lu et al.VLDB 2026
- Swift: Fast Performance Tuning with GAN-Generated ConfigurationsChao Chen, Shixin Huang, Xuehai Qian, Zhibin YuUSENIX ATC 2025 ยท 1 citation
- Towards Dynamic and Safe Configuration Tuning for Cloud DatabasesXinyi Zhang, Hong Wu, Yang Li, Jian Tan et al.SIGMOD 2022 ยท 62 citations
- MCTuner: Spatial Decomposition-Enhanced Database Tuning via LLM-Guided ExplorationZihan Yan, Rui Xi, Mengshu HouSIGMOD 2026 ยท 3 citations
