LOCAT: Low-Overhead Online Configuration Auto-Tuning of Spark SQL Applications
Jinhan Xin, Kai Hwang, Zhibin Yu
摘要
Spark SQL has been widely deployed in industry but it is challenging to tune its performance. Recent studies try to employ machine learning (ML) to solve this problem. They however suffer from two drawbacks. First, it takes a long time (high overhead) to collect training samples. Second, the optimal configuration for one input data size of the same application might not be optimal for others.
To address these issues, we propose a novel Bayesian Optimization (BO) based approach named LOCAT to automatically tune the configurations of Spark SQL applications online. LOCAT innovates three techniques. The first technique, named QCSA, eliminates the configuration-insensitive queries by Query Configuration Sensitivity Analysis (QCSA) when collecting training samples. The second technique, dubbed DAGP, is a Datasize-Aware Gaussian Process (DAGP) which models the performance of an application as a distribution of functions of configuration parameters as well as input data size. The third technique, called IICP, Identifies Important Configuration Parameters (IICP) with respect to performance and only tunes the important parameters. As such, LOCAT can tune the configurations of a Spark SQL application with low overhead and adapt to different input data sizes.
We employ Spark SQL applications from benchmark suites𝑇 𝑃𝐶-𝐷𝑆, 𝑇 𝑃𝐶 -𝐻 , and 𝐻𝑖𝐵𝑒𝑛𝑐ℎ running on two significantly different clusters, a four-node ARM cluster and an eight-node x86 cluster, to evaluate LOCAT. The experimental results on the ARM cluster show that LOCAT accelerates the optimization procedures of Tuneful [22], DAC [66], , and QTune [37] by factors of 6.4×, 7.0×, 4.1×, and 9.7× on average, respectively. On the x86 cluster, LOCAT reduces the optimization time of Tuneful, DAC, GBO-RL, and QTune by factors of 6.4×, 6.3×, 4.0×, and 9.2× on average, respectively. Moreover, LOCAT improves the performance of the applications on
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- OPPerTune: Post-Deployment Configuration Tuning of Services Made EasyGagan Somashekar, Karan Tandon, Anush Kini, Chieh-Chun Chang 等NSDI 2024 · 被引用 24 次
- VDTuner: Automated Performance Tuning for Vector Data Management SystemsTiannuo Yang, Wen Hu, Wangqi Peng, Yusen Li 等ICDE 2024 · 被引用 12 次
- A Spark Optimizer for Adaptive, Fine-Grained Parameter TuningChenghao Lyu, Qi Fan, Philippe Guyard, Yanlei DiaoVLDB 2024 · 被引用 9 次
- LEAP: A Low-cost Spark SQL Query Optimizer using Pairwise ComparisonJunhao Ye, Jiahui Li, Lu Chen, Yuren Mao 等VLDB 2025 · 被引用 2 次
- InferLog: Accelerating LLM Inference for Online Log Parsing via ICL-oriented Prefix CachingYilun Wang, Pengfei Chen, Haiyu Huang, Zilong He 等ICSE 2026 · 被引用 1 次
它引用的顶会 Paper1
相关 Paper
- Adaptive Code Learning for Spark Configuration TuningChen Lin, Junqing Zhuang, Jiadong Feng, Hui Li 等ICDE 2022 · 被引用 28 次
- MFTune: An Efficient Multi-fidelity Framework for Spark SQL Configuration TuningBeicheng Xu, Lingching Tung, Yuchen Wang, Yupeng Lu 等VLDB 2026
- Swift: Fast Performance Tuning with GAN-Generated ConfigurationsChao Chen, Shixin Huang, Xuehai Qian, Zhibin YuUSENIX ATC 2025 · 被引用 1 次
- Towards Dynamic and Safe Configuration Tuning for Cloud DatabasesXinyi Zhang, Hong Wu, Yang Li, Jian Tan 等SIGMOD 2022 · 被引用 62 次
- MCTuner: Spatial Decomposition-Enhanced Database Tuning via LLM-Guided ExplorationZihan Yan, Rui Xi, Mengshu HouSIGMOD 2026 · 被引用 3 次
