Fine-Grained Modeling and Optimization for Intelligent Resource Management in Big Data Processing
Chenghao Lyu, Qi Fan, Fei Song, Arnab Sinha, Yanlei Diao, Wei Chen, Li Ma, Yihui Feng, Yaliang Li, Kai Zeng, Jingren Zhou
摘要
Big data processing at the production scale presents a highly complex environment for resource optimization (RO), a problem crucial for meeting performance goals and budgetary constraints of analytical users. The RO problem is challenging because it involves a set of decisions (the partition count, placement of parallel instances on machines, and resource allocation to each instance), requires multi-objective optimization (MOO), and is compounded by the scale and complexity of big data systems while having to meet stringent time constraints for scheduling. This paper presents a MaxCompute based integrated system to support multi-objective resource optimization via fine-grained instance-level modeling and optimization. We propose a new architecture that breaks RO into a series of simpler problems, new fine-grained predictive models, and novel optimization methods that exploit these models to make effective instance-level RO decisions well under a second. Evaluation using production workloads shows that our new RO system could reduce 37--72% latency and 43--78% cost at the same time, compared to the current optimizer and scheduler, while running in 0.02-0.23s.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- The Holon Approach for Simultaneously Tuning Multiple Components in a Self-Driving Database Management System with Machine Learning via Synthesized Proto-ActionsWilliam Zhang, Wan Shen Lim, Matthew Butrovich, Andrew PavloVLDB 2024 · 被引用 13 次
- A Spark Optimizer for Adaptive, Fine-Grained Parameter TuningChenghao Lyu, Qi Fan, Philippe Guyard, Yanlei DiaoVLDB 2024 · 被引用 9 次
- T3: Accurate and Fast Performance Prediction for Relational Database Systems With Compiled Decision TreesMaximilian Rieger, Thomas NeumannSIGMOD 2025 · 被引用 5 次
- Improving DBMS Scheduling Decisions with Accurate Performance Prediction on Concurrent QueriesZiniu Wu, Markos Markakis, Chunwei Liu, Peter Baile Chen 等VLDB 2025
- Graph Transformers for Query Plan Representation: Potentials and ChallengesChenghao Lyu, Guillaume Lachaud, Gabriel Lozano, Yanlei DiaoVLDB 2025
它引用的顶会 Paper21
- An End-to-End Learning-based Cost EstimatorJi Sun, Guoliang LiVLDB 2020 · 被引用 251 次
- Deep Unsupervised Cardinality EstimationZongheng Yang, Eric Liang, Amog Kamsetty, Chenggang Wu 等VLDB 2020 · 被引用 206 次
- NeuroCard: One Cardinality Estimator for All TablesZongheng Yang, Amog Kamsetty, Sifei Luan, Eric Liang 等VLDB 2021 · 被引用 138 次
- FLAT: Fast, Lightweight and Accurate Method for Cardinality EstimationRong Zhu, Ziniu Wu, Yuxing Han, Kai Zeng 等VLDB 2021 · 被引用 120 次
- ResTune: Resource Oriented Tuning Boosted by Meta-Learning for Cloud DatabasesXinyi Zhang, Hong Wu, Zhuo Chang, Shuowei Jin 等SIGMOD 2021 · 被引用 113 次
相关 Paper
- Spark-based Cloud Data Analytics using Multi-Objective OptimizationFei Song, Khaled Zaouk, Chenghao Lyu, Arnab Sinha 等ICDE 2021 · 被引用 15 次
- ReLoca: Optimize Resource Allocation for Data-parallel Jobs using Deep LearningZhiyao Hu, Dongsheng Li, Dongxiang Zhang, Yixin ChenINFOCOM 2020 · 被引用 3 次
- Cost Models for Big Data Query Processing: Learning, Retrofitting, and Our FindingsTarique Siddiqui, Alekh Jindal, Shi Qiao, Hiren Patel 等SIGMOD 2020 · 被引用 80 次
- A House United Within Itself: SLO-Awareness for On-Premises Containerized ML Inference Clusters via FaroBeomyeol Jeon, Chen Wang, Diana Arroyo, Alaa Youssef 等EuroSys 2025
- CLITE: Efficient and QoS-Aware Co-Location of Multiple Latency-Critical Jobs for Warehouse Scale ComputersTirthak Patel, Devesh TiwariHPCA 2020 · 被引用 153 次
