Sluice: End-to-End Latency Guarantee for Long-running Dataflow Systems
Zhaochen She, Yancan Mao, Richard T. B. Ma
Abstract
End-to-end latency is a key performance metric for long-running dataflow systems, as it directly impacts the timeliness and quality of results in real-time applications. In cloud environments, dynamic resource scaling is essential for handling workload fluctuations efficiently. However, maintaining consistent latency remains challenging due to resource contention and scaling-induced disruptions. Existing scaling techniques improve efficiency but fall short in guaranteeing latency, as they lack accurate latency estimation and timely scaling decisions. We present Sluice, a general framework for guaranteeing end-to-end latency. The core insight behind Sluice is to decompose latency into intrinsic and extrinsic components, isolating predictable, controllable delays from external disruptions like scaling and scheduling. Sluice dynamically reserves a buffer for extrinsic latency and scales resources to tightly control intrinsic latency. At its core, Sluice integrates (1) a novel Latency Estimation Model (LEM) that uses instantaneous backlog sizes along the critical path, and (2) a scaling controller that leverages LEM to make latency-aware, resource-efficient scaling decisions. We implement Sluice on Apache Flink and evaluate it using diverse real-world workloads. Results show that Sluice consistently enforces end-to-end latency guarantees under dynamic conditions while maintaining competitive resource efficiency, outperforming existing approaches in both latency control and task usage.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 53acd222-df38-4a4d-b963-d8cd57376537Related papers
- StreamSwitch: Fulfilling Latency Service-Layer Agreement for Stateful StreamingZhaochen She, Yancan Mao, Hailin Xiang, Xin Wang et al.INFOCOM 2023 · 5 citations
- Towards Fine-Grained Scalability for Stateful Stream Processing SystemsYunfan Qing, Wenli ZhengICDE 2025 · 2 citations
- Latency-Oriented Elastic Memory Management at Task-Granularity for Stateful Streaming ProcessingRengan Dou, Richard T. B. MaINFOCOM 2023 · 2 citations
- On Modular Learning of Distributed Systems for Predicting End-to-End LatencyChieh-Jan Mike Liang, Zilin Fang, Yuqing Xie, Fan Yang et al.NSDI 2023 · 19 citations
- Fugue: Online Elasticity for Distributed Stateful Stream ProcessingYuqiu Zhang, Yunhao Mao, Hans-Arno JacobsenVLDB 2026
