Adaptive Code Learning for Spark Configuration Tuning
Chen Lin, Junqing Zhuang, Jiadong Feng, Hui Li, Xuanhe Zhou, Guoliang Li
摘要
Configuration tuning is vital to optimize the performance of big data analysis platforms like Spark. Existing methods (e.g. auto-tuning relational databases) are not effective for tuning Spark, because the unique characteristics of Spark pose new challenges to configuration tuning. (C1) The Spark applications own various code structures and semantics, and the code features significantly affect Spark performance and configuration selection; (C2) Spark applications are extremely time-consuming on big data. It is infeasible for approaches such as Bayesian Optimization and Reinforcement Learning to collect sufficient training instances or repeatedly execute the applications; (C3) Spark supports various analytical applications and the tuning system needs to adapt to different applications. To address these challenges, we propose a LIghtweighT knob rEcommender system (LITE) for auto-tuning Spark configurations on various analytical applications and large-scale datasets. We first propose a code learning framework that can utilize code features to learn complex correlations between application performance and knob values (addressing C1). We then propose a lightweight auto-tuning method that migrates the knowledge learned from small-scale datasets to large-scale datasets (addressing C2). Next, to generalize to different Spark applications, we propose an adaptive model update approach to fine-tune the model via adversarial learning with newly collected feedback (addressing C3). Extensive experiments showed that LITE achieves much better performance compared with state-of-the-art auto-tuning methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- OPPerTune: Post-Deployment Configuration Tuning of Services Made EasyGagan Somashekar, Karan Tandon, Anush Kini, Chieh-Chun Chang 等NSDI 2024 · 被引用 24 次
- VDTuner: Automated Performance Tuning for Vector Data Management SystemsTiannuo Yang, Wen Hu, Wangqi Peng, Yusen Li 等ICDE 2024 · 被引用 12 次
- A Spark Optimizer for Adaptive, Fine-Grained Parameter TuningChenghao Lyu, Qi Fan, Philippe Guyard, Yanlei DiaoVLDB 2024 · 被引用 9 次
- LST-Bench: Benchmarking Log-Structured Tables in the CloudJesús Camacho-Rodríguez, Ashvin Agrawal, Anja Gruenheid, Ashit Gosalia 等SIGMOD 2024 · 被引用 7 次
- AgentTune: An Agent-Based Large Language Model Framework for Database Knob TuningYiyan Li, Haoyang Li, Jing Zhang, Renata Borovica-Gajic 等SIGMOD 2026 · 被引用 5 次
它引用的顶会 Paper8
- An End-to-End Learning-based Cost EstimatorJi Sun, Guoliang LiVLDB 2020 · 被引用 251 次
- Reinforcement Learning with Tree-LSTM for Join Order SelectionXiang Yu, Guoliang Li, Chengliang Chai, Nan TangICDE 2020 · 被引用 168 次
- ResTune: Resource Oriented Tuning Boosted by Meta-Learning for Cloud DatabasesXinyi Zhang, Hong Wu, Zhuo Chang, Shuowei Jin 等SIGMOD 2021 · 被引用 113 次
- An Inquiry into Machine Learning-based Automatic Configuration Tuning Services on Real-World Database Management SystemsDana Van Aken, Dongsheng Yang, Sebastien Brillard, Ari Fiorino 等VLDB 2021 · 被引用 108 次
- Query Performance Prediction for Concurrent Queries using Graph EmbeddingXuanhe Zhou, Ji Sun, Guoliang Li, Jianhua FengVLDB 2020 · 被引用 96 次
相关 Paper
- LOCAT: Low-Overhead Online Configuration Auto-Tuning of Spark SQL ApplicationsJinhan Xin, Kai Hwang, Zhibin YuSIGMOD 2022 · 被引用 34 次
- MFTune: An Efficient Multi-fidelity Framework for Spark SQL Configuration TuningBeicheng Xu, Lingching Tung, Yuchen Wang, Yupeng Lu 等VLDB 2026
- Scarf: Self-Adaptive Tuning via Multi-Objective Reinforcement Learning for Apache FlinkLiu Liu, Shenghao Gong, Ziquan Fang, Yunjun GaoVLDB 2026
- Swift: Fast Performance Tuning with GAN-Generated ConfigurationsChao Chen, Shixin Huang, Xuehai Qian, Zhibin YuUSENIX ATC 2025 · 被引用 1 次
- Why Database Manuals Are Not Enough: Efficient and Reliable Configuration Tuning for DBMSs via Code-Driven LLM AgentsXinyi Zhang, Tiantian Chen, Zhentao Han, Zhaoyan Hong 等VLDB 2026 · 被引用 5 次
