Learning from the Past: Adaptive Parallelism Tuning for Stream Processing Systems
Yuxing Han, Lixiang Chen, Haoyu Wang, Zhanghao Chen, Yifan Zhang, Chengcheng Yang, Kongzhang Hao, Zhengyi Yang
摘要
Distributed stream processing systems rely on the dataflow model to define and execute streaming jobs, organizing computations as Directed Acyclic Graphs (DAGs) of operators. Adjusting the parallelism of these operators is crucial to handling fluctuating workloads efficiently while balancing resource usage and processing performance. However, existing methods often fail to effectively utilize execution histories or fully exploit DAG structures, limiting their ability to identify bottlenecks and determine the optimal parallelism. In this paper, we propose StreamTune, a novel approach for adaptive parallelism tuning in stream processing systems. StreamTune incorporates a pre-training and fine-tuning framework that leverages global knowledge from historical execution data for job-specific parallelism tuning. In the pre-training phase, StreamTune clusters the historical data with Graph Edit Distance and pre-trains a Graph Neural Network-based encoder per cluster to capture the correlation between the operator parallelism, DAG structures, and the identified operator-level bottlenecks. In the online tuning phase, Stream-Tu ne iteratively refines operator parallelism recommendations using an operator-level bottleneck prediction model enforced with a monotonic constraint, which aligns with the observed system performance behavior. Evaluation results demonstrate that StreamTune reduces reconfigurations by up to 29.6% and parallelism degrees by up to 30.8% in Apache Flink under a synthetic workload. In Timely Dataflow, StreamTune achieves up to an 83.3% reduction in parallelism degrees while maintaining comparable processing performance under the Nexmark benchmark, when compared to the state-of-the-art methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper9
- EvolveGCN: Evolving Graph Convolutional Networks for Dynamic GraphsAldo Pareja, Giacomo Domeniconi, Jie Chen, Tengfei Ma 等AAAI 2020 · 被引用 1,429 次
- Certified Monotonic Neural NetworksXingchao Liu, Xing Han, Na Zhang, Qiang LiuNeurIPS 2020 · 被引用 116 次
- ResTune: Resource Oriented Tuning Boosted by Meta-Learning for Cloud DatabasesXinyi Zhang, Hong Wu, Zhuo Chang, Shuowei Jin 等SIGMOD 2021 · 被引用 113 次
- LlamaTune: Sample-Efficient DBMS Configuration TuningKonstantinos Kanellis, Cong Ding, Brian Kroth, Andreas Müller 等VLDB 2022 · 被引用 73 次
- Constrained Monotonic Neural NetworksDavor Runje, Sharath M. ShankaranarayanaICML 2023 · 被引用 61 次
相关 Paper
- ZERoTuNE: Learned Zero-Shot Cost Models for Parallelism Tuning in Stream ProcessingPratyush Agnihotri, Boris Koldehofe, Paul Stiegele, Roman Heinrich 等ICDE 2024 · 被引用 13 次
- ContTune: Continuous Tuning by Conservative Bayesian Optimization for Distributed Stream Data Processing SystemsJinqing Lian, Xinyi Zhang, Yingxia Shao, Zenglin Pu 等VLDB 2023 · 被引用 8 次
- Stream processing with dependency-guided synchronizationKonstantinos Kallas, Filip Niksic, Caleb Stanford, Rajeev AlurPPoPP 2022 · 被引用 4 次
- CrystalPerf: Learning to Characterize the Performance of Dataflow Computation through Code AnalysisHuangshi Tian, Minchen Yu, Wei WangUSENIX ATC 2021 · 被引用 2 次
- SaSPartitioner: A Self-Adaptive Streaming Partitioner Using Deep Reinforcement LearningShenghao Gong, Liu Liu, Ziquan Fang, Yunjun Gao 等ICDE 2026
