SaSPartitioner: A Self-Adaptive Streaming Partitioner Using Deep Reinforcement Learning
Shenghao Gong, Liu Liu, Ziquan Fang, Yunjun Gao, Yaofeng Tu
摘要
Stream processing has been widely adopted in online services, such as e-commerce platforms and real-time query systems, due to its ability to leverage distributed parallel computing for low-latency, large-scale data processing. However, data skew, which results in imbalanced load distribution across computational nodes, can significantly increase processing latency and degrade resource efficiency. Most existing stream processing systems adopt partitioning strategies based on data volume, estimating system load using the volume of data in each partition. As a result, they often struggle to adapt promptly to dynamic changes in data distribution caused by streaming workloads. To this end, we propose SaSPartitioner, a self-adaptive partitioning framework based on system-metric modeling that leverages deep reinforcement learning (DRL) and real-time system metrics. Specifically, we formulate the streaming partitioning problem as a Markov decision process (MDP), collecting real-time metrics such as CPU and memory usage of streaming operators. Besides, we design both offline and online reinforcement learning models to enable seamless adaptation to dynamic shifts in data distribution. To further enhance efficiency, we introduce two acceleration methods, i.e., partition masking and hot key selection, to prune the state and action space of the DRL model. Extensive experiments on Apache Flink show that SaSPartitioner significantly outperforms existing stream partitioning methods, achieving up to a increase in throughput while maintaining robustness under rapidly changing data distributions.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Dalton: Learned Partitioning for Distributed Data StreamsEleni Zapridou, Ioannis Mytilinis, Anastasia AilamakiVLDB 2023 · 被引用 25 次
- SASPAR: Shared Adaptive Stream PartitioningJeyhun Karimov, Hans-Arno JacobsenICDE 2023 · 被引用 3 次
- Scarf: Self-Adaptive Tuning via Multi-Objective Reinforcement Learning for Apache FlinkLiu Liu, Shenghao Gong, Ziquan Fang, Yunjun GaoVLDB 2026
- Generalizable Resource Allocation in Stream Processing via Deep Reinforcement LearningXiang Ni, Jing Li, Mo Yu, Wang Zhou 等AAAI 2020 · 被引用 24 次
- Learning from the Past: Adaptive Parallelism Tuning for Stream Processing SystemsYuxing Han, Lixiang Chen, Haoyu Wang, Zhanghao Chen 等ICDE 2025 · 被引用 2 次
