Spark-based Cloud Data Analytics using Multi-Objective Optimization
Fei Song, Khaled Zaouk, Chenghao Lyu, Arnab Sinha, Qi Fan, Yanlei Diao, Prashant J. Shenoy
摘要
Data analytics in the cloud has become an integral part of enterprise businesses. Big data analytics systems, however, still lack the ability to take task objectives such as user performance goals and budgetary constraints and automatically configure an analytic job to achieve these objectives. This paper presents UDAO, a Spark-based Unified Data Analytics Optimizer that can automatically determine a cluster configuration with a suitable number of cores as well as other system parameters that best meet the task objectives. At a core of our work is a principled multi-objective optimization (MOO) approach that computes a Pareto optimal set of configurations to reveal tradeoffs between different objectives, recommends a new Spark configuration that best explores such tradeoffs, and employs novel optimizations to enable such recommendations within a few seconds. Detailed experiments using benchmark workloads show that our MOO techniques provide a 2-50× speedup over existing MOO methods, while offering good coverage of the Pareto frontier. Compared to Ottertune, a state-of-the-art performance tuning system, UDAO recommends Spark configurations that yield 26%-49% reduction of running time of the TPCx-BB benchmark while adapting to different user preferences on multiple objectives.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Fine-Grained Modeling and Optimization for Intelligent Resource Management in Big Data ProcessingChenghao Lyu, Qi Fan, Fei Song, Arnab Sinha 等VLDB 2022 · 被引用 14 次
- A Spark Optimizer for Adaptive, Fine-Grained Parameter TuningChenghao Lyu, Qi Fan, Philippe Guyard, Yanlei DiaoVLDB 2024 · 被引用 9 次
- Scarf: Self-Adaptive Tuning via Multi-Objective Reinforcement Learning for Apache FlinkLiu Liu, Shenghao Gong, Ziquan Fang, Yunjun GaoVLDB 2026
- Graph Transformers for Query Plan Representation: Potentials and ChallengesChenghao Lyu, Guillaume Lachaud, Gabriel Lozano, Yanlei DiaoVLDB 2025
它引用的顶会 Paper3
- Differentiable Expected Hypervolume Improvement for Parallel Multi-Objective Bayesian OptimizationSamuel Daulton, Maximilian Balandat, Eytan BakshyNeurIPS 2020 · 被引用 428 次
- An End-to-End Learning-based Cost EstimatorJi Sun, Guoliang LiVLDB 2020 · 被引用 251 次
- Cost Models for Big Data Query Processing: Learning, Retrofitting, and Our FindingsTarique Siddiqui, Alekh Jindal, Shi Qiao, Hiren Patel 等SIGMOD 2020 · 被引用 80 次
相关 Paper
- Adaptive Code Learning for Spark Configuration TuningChen Lin, Junqing Zhuang, Jiadong Feng, Hui Li 等ICDE 2022 · 被引用 28 次
- Swift: Fast Performance Tuning with GAN-Generated ConfigurationsChao Chen, Shixin Huang, Xuehai Qian, Zhibin YuUSENIX ATC 2025 · 被引用 1 次
- MONSOON: Multi-Step Optimization and Execution of Queries with Partially Obscured PredicatesSourav Sikdar, Chris JermaineSIGMOD 2020 · 被引用 6 次
- An Inquiry into Machine Learning-based Automatic Configuration Tuning Services on Real-World Database Management SystemsDana Van Aken, Dongsheng Yang, Sebastien Brillard, Ari Fiorino 等VLDB 2021 · 被引用 108 次
- QaaD (Query-as-a-Data): Scalable Execution of Massive Number of Small Queries in SparkYeonsu Park, Byungchul Tak, Wook-Shin HanSIGMOD 2023 · 被引用 3 次
