Adaptive Code Learning for Spark Configuration Tuning
Chen Lin, Junqing Zhuang, Jiadong Feng, Hui Li, Xuanhe Zhou, Guoliang Li
Abstract
Configuration tuning is vital to optimize the performance of big data analysis platforms like Spark. Existing methods (e.g. auto-tuning relational databases) are not effective for tuning Spark, because the unique characteristics of Spark pose new challenges to configuration tuning. (C1) The Spark applications own various code structures and semantics, and the code features significantly affect Spark performance and configuration selection; (C2) Spark applications are extremely time-consuming on big data. It is infeasible for approaches such as Bayesian Optimization and Reinforcement Learning to collect sufficient training instances or repeatedly execute the applications; (C3) Spark supports various analytical applications and the tuning system needs to adapt to different applications. To address these challenges, we propose a LIghtweighT knob rEcommender system (LITE) for auto-tuning Spark configurations on various analytical applications and large-scale datasets. We first propose a code learning framework that can utilize code features to learn complex correlations between application performance and knob values (addressing C1). We then propose a lightweight auto-tuning method that migrates the knowledge learned from small-scale datasets to large-scale datasets (addressing C2). Next, to generalize to different Spark applications, we propose an adaptive model update approach to fine-tune the model via adversarial learning with newly collected feedback (addressing C3). Extensive experiments showed that LITE achieves much better performance compared with state-of-the-art auto-tuning methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 06f2aca7-d4ce-4965-8557-c443a1e276e6Cited by top-tier papers7
- OPPerTune: Post-Deployment Configuration Tuning of Services Made EasyGagan Somashekar, Karan Tandon, Anush Kini, Chieh-Chun Chang et al.NSDI 2024 · 24 citations
- VDTuner: Automated Performance Tuning for Vector Data Management SystemsTiannuo Yang, Wen Hu, Wangqi Peng, Yusen Li et al.ICDE 2024 · 12 citations
- A Spark Optimizer for Adaptive, Fine-Grained Parameter TuningChenghao Lyu, Qi Fan, Philippe Guyard, Yanlei DiaoVLDB 2024 · 9 citations
- LST-Bench: Benchmarking Log-Structured Tables in the CloudJesús Camacho-Rodríguez, Ashvin Agrawal, Anja Gruenheid, Ashit Gosalia et al.SIGMOD 2024 · 7 citations
- AgentTune: An Agent-Based Large Language Model Framework for Database Knob TuningYiyan Li, Haoyang Li, Jing Zhang, Renata Borovica-Gajic et al.SIGMOD 2026 · 5 citations
Builds on8
- An End-to-End Learning-based Cost EstimatorJi Sun, Guoliang LiVLDB 2020 · 251 citations
- Reinforcement Learning with Tree-LSTM for Join Order SelectionXiang Yu, Guoliang Li, Chengliang Chai, Nan TangICDE 2020 · 168 citations
- ResTune: Resource Oriented Tuning Boosted by Meta-Learning for Cloud DatabasesXinyi Zhang, Hong Wu, Zhuo Chang, Shuowei Jin et al.SIGMOD 2021 · 113 citations
- An Inquiry into Machine Learning-based Automatic Configuration Tuning Services on Real-World Database Management SystemsDana Van Aken, Dongsheng Yang, Sebastien Brillard, Ari Fiorino et al.VLDB 2021 · 108 citations
- Query Performance Prediction for Concurrent Queries using Graph EmbeddingXuanhe Zhou, Ji Sun, Guoliang Li, Jianhua FengVLDB 2020 · 96 citations
Related papers
- LOCAT: Low-Overhead Online Configuration Auto-Tuning of Spark SQL ApplicationsJinhan Xin, Kai Hwang, Zhibin YuSIGMOD 2022 · 34 citations
- MFTune: An Efficient Multi-fidelity Framework for Spark SQL Configuration TuningBeicheng Xu, Lingching Tung, Yuchen Wang, Yupeng Lu et al.VLDB 2026
- Scarf: Self-Adaptive Tuning via Multi-Objective Reinforcement Learning for Apache FlinkLiu Liu, Shenghao Gong, Ziquan Fang, Yunjun GaoVLDB 2026
- Swift: Fast Performance Tuning with GAN-Generated ConfigurationsChao Chen, Shixin Huang, Xuehai Qian, Zhibin YuUSENIX ATC 2025 · 1 citation
- Why Database Manuals Are Not Enough: Efficient and Reliable Configuration Tuning for DBMSs via Code-Driven LLM AgentsXinyi Zhang, Tiantian Chen, Zhentao Han, Zhaoyan Hong et al.VLDB 2026 · 5 citations
