A Spark Optimizer for Adaptive, Fine-Grained Parameter Tuning
Chenghao Lyu, Qi Fan, Philippe Guyard, Yanlei Diao
摘要
As Spark becomes a common big data analytics platform, its growing complexity makes automatic tuning of numerous parameters critical for performance. Our work on Spark parameter tuning is particularly motivated by two recent trends: Spark's Adaptive Query Execution (AQE) based on runtime statistics, and the increasingly popular Spark cloud deployments that make cost-performance reasoning crucial for the end user. This paper presents our design of a Spark optimizer that controls all tunable parameters of each query in the new AQE architecture to explore its performance benefits and, at the same time, casts the tuning problem in the theoretically sound multi-objective optimization (MOO) setting to better adapt to user cost-performance preferences. To this end, we propose a novel hybrid compile-time/runtime approach to multi-granularity tuning of diverse, correlated Spark parameters, as well as a suite of modeling and optimization techniques to solve the tuning problem in the MOO setting while meeting the stringent time constraint of 1--2 seconds for cloud use. Evaluation results using TPC-H and TPC-DS benchmarks demonstrate the superior performance of our approach: (i ) When prioritizing latency, it achieves 63% and 65% reduction for TPC-H and TPC-DS, respectively, under an average solving time of 0.7--0.8 sec, outperforming the most competitive MOO method that reduces only 18--25% latency with 2.6--15 sec solving time. (ii) When shifting preferences between latency and cost, our approach dominates the solutions of alternative methods, exhibiting superior adaptability to varying preferences.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Scarf: Self-Adaptive Tuning via Multi-Objective Reinforcement Learning for Apache FlinkLiu Liu, Shenghao Gong, Ziquan Fang, Yunjun GaoVLDB 2026
- Graph Transformers for Query Plan Representation: Potentials and ChallengesChenghao Lyu, Guillaume Lachaud, Gabriel Lozano, Yanlei DiaoVLDB 2025
- LakeHelm: Zero-Shot Lakehouse Advisor for Joint Engine-Format Selection and ConfigurationZhongwei Xu, Siyuan Dong, Haotian Gong, Donna Pham 等VLDB 2026
它引用的顶会 Paper20
- Differentiable Expected Hypervolume Improvement for Parallel Multi-Objective Bayesian OptimizationSamuel Daulton, Maximilian Balandat, Eytan BakshyNeurIPS 2020 · 被引用 428 次
- FLAT: Fast, Lightweight and Accurate Method for Cardinality EstimationRong Zhu, Ziniu Wu, Yuxing Han, Kai Zeng 等VLDB 2021 · 被引用 120 次
- ResTune: Resource Oriented Tuning Boosted by Meta-Learning for Cloud DatabasesXinyi Zhang, Hong Wu, Zhuo Chang, Shuowei Jin 等SIGMOD 2021 · 被引用 113 次
- Flow-Loss: Learning Cardinality Estimates That MatterParimarjan Negi, Ryan Marcus, Andreas Kipf, Hongzi Mao 等VLDB 2021 · 被引用 102 次
- Deep Learning Models for Selectivity Estimation of Multi-Attribute QueriesShohedul Hasan, Saravanan Thirumuruganathan, Jees Augustine, Nick Koudas 等SIGMOD 2020 · 被引用 101 次
相关 Paper
- Spark-based Cloud Data Analytics using Multi-Objective OptimizationFei Song, Khaled Zaouk, Chenghao Lyu, Arnab Sinha 等ICDE 2021 · 被引用 15 次
- MFTune: An Efficient Multi-fidelity Framework for Spark SQL Configuration TuningBeicheng Xu, Lingching Tung, Yuchen Wang, Yupeng Lu 等VLDB 2026
- LOCAT: Low-Overhead Online Configuration Auto-Tuning of Spark SQL ApplicationsJinhan Xin, Kai Hwang, Zhibin YuSIGMOD 2022 · 被引用 34 次
- Fine-Grained Modeling and Optimization for Intelligent Resource Management in Big Data ProcessingChenghao Lyu, Qi Fan, Fei Song, Arnab Sinha 等VLDB 2022 · 被引用 14 次
- AQETuner: Reliable Query-level Configuration Tuning for Analytical Query EnginesLixiang Chen, Yuxing Han, Yu Chen, Xing Chen 等VLDB 2025 · 被引用 4 次
