Sponge: Fast Reactive Scaling for Stream Processing with Serverless Frameworks
Won Wook Song, Taegeon Um, Sameh Elnikety, Myeongjae Jeon, Byung-Gon Chun
Abstract
Streaming workloads deal with data that is generated in realtime. This data is often unpredictable and changes rapidly in volume. To deal with these fluctuations, current systems aim to dynamically scale in and out, redistribute, and migrate computing tasks across a cluster of machines. While many prior works have focused on reducing the overhead of system reconfiguration and state migration on pre-allocated cluster resources, these approaches still face significant challenges in meeting latency SLOs at low operational costs, especially upon facing unpredictable bursty loads.
In this paper, we propose Sponge, a new stream processing system that enables fast reactive scaling of long-running stream queries by leveraging serverless framework (SF) instances. Sponge absorbs sudden, unpredictable increases in input loads from existing VMs with low latency and cost by taking advantage of the fact that SF instances can be initiated quickly, in just a few hundred milliseconds. Sponge efficiently tracks a small number of metrics to quickly detect bursty loads and make fast scaling decisions based on these metrics. Moreover, by incorporating optimization logic at compile-time and triggering fast data redirection and partial-state merging mechanisms at runtime, Sponge avoids optimization and state migration overheads during runtime while efficiently offloading bursty loads from existing VMs to new SF instances. Our evaluation on AWS EC2 and Lambda using the NEXMark benchmark shows that Sponge promptly reacts to bursty input loads, reducing 99 th -percentile tail latencies by 88% on average compared to other stream query scaling methods on VMs. Sponge also reduces cost by 83% compared to methods that over-provision VMs to handle unpredictable bursty loads.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a9a95bdf-00b0-4163-8782-3deba1e54a85Cited by top-tier papers4
- ESG: Pipeline-Conscious Efficient Scheduling of DNN Workflows on Serverless Platforms with Shareable GPUsXinning Hui, Yuanchao Xu, Zhishan Guo, Xipeng ShenHPDC 2024 · 11 citations
- Burst Computing: Quick, Sudden, Massively Parallel Processing on Serverless ResourcesDaniel Barcelona Pons, Aitor Arjona, Pedro García López, Enrique Molina-Giménez et al.USENIX ATC 2025 · 3 citations
- Towards Fine-Grained Scalability for Stateful Stream Processing SystemsYunfan Qing, Wenli ZhengICDE 2025 · 2 citations
- Blaze: Holistic Caching for Iterative Data ProcessingWon Wook Song, Jeongyoon Eo, Taegeon Um, Myeongjae Jeon et al.EuroSys 2024 · 1 citation
Builds on3
- Lambada: Interactive Data Analytics on Cold Data Using Serverless Cloud InfrastructureIngo Müller, Renato Marroquín, Gustavo AlonsoSIGMOD 2020 · 135 citations
- Rhino: Efficient Management of Very Large Distributed State for Stream Processing EnginesBonaventura Del Monte, Steffen Zeuch, Tilmann Rabl, Volker MarklSIGMOD 2020 · 56 citations
- Meces: Latency-efficient Rescaling via Prioritized State Migration for Stateful Distributed Stream Processing SystemsRong Gu, Han Yin, Weichang Zhong, Chunfeng Yuan et al.USENIX ATC 2022 · 22 citations
Related papers
- Batch: machine learning inference serving on serverless platforms with adaptive batchingAhsan Ali, Riccardo Pinciroli, Feng Yan, Evgenia SmirniSC 2020 · 184 citations
- StreamBox: A Lightweight GPU SandBox for Serverless Inference WorkflowHao Wu, Yue Yu, Junxiao Deng, Shadi Ibrahim et al.USENIX ATC 2024 · 21 citations
- StreamSwitch: Fulfilling Latency Service-Layer Agreement for Stateful StreamingZhaochen She, Yancan Mao, Hailin Xiang, Xin Wang et al.INFOCOM 2023 · 5 citations
- SASPAR: Shared Adaptive Stream PartitioningJeyhun Karimov, Hans-Arno JacobsenICDE 2023 · 3 citations
- Emma: Elastic Multi-Resource Management for Realtime Stream ProcessingRengan Dou, Xin Wang, Richard T. B. MaINFOCOM 2024 · 2 citations
