Caerus: NIMBLE Task Scheduling for Serverless Analytics
Hong Zhang, Yupeng Tang, Anurag Khandelwal, Jingrong Chen, Ion Stoica
Abstract
Serverless platforms facilitate transparent resource elasticity and fine-grained billing, making them an attractive choice for data analytics. We find that while server-centric analytics frameworks typically optimize for job completion time (JCT), resource utilization and isolation via inter-job scheduling policies, serverless analytics requires optimizing for JCT and cost of execution instead, introducing a new scheduling problem. We present Caerus, a task scheduler for serverless analytics frameworks that employs a fine-grained NIMBLE scheduling algorithm to solve this problem. NIMBLE efficiently pipelines task executions within a job, minimizing execution cost while being Pareto-optimal between cost and JCT for arbitrary analytics jobs. To this end, NIMBLE models a wide range of execution parameters -pipelineable and non-piplineable data dependencies, data generation, consumption and processing rates, etc. -to determine the ideal task launch times. Our evaluation results show that in practice, Caerus is able to achieve both optimal cost and JCT for queries across a wide range of analytics workloads.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9623202b-b376-4434-a714-d4926d6ee6c2Cited by top-tier papers12
- Fast Vector Query Processing for Large Datasets Beyond GPU Memory with Reordered PipeliningZili Zhang, Fangyue Liu, Gang Huang, Xuanzhe Liu et al.NSDI 2024 · 37 citations
- Demeter: Fine-grained Function Orchestration for Geo-distributed Serverless AnalyticsXiaofei Yue, Song Yang, Liehuang Zhu, Stojan Trajanovski et al.INFOCOM 2024 · 13 citations
- Making Serverless Pay-For-Use a Reality with LeopardTingjia Cao, Andrea C. Arpaci-Dusseau, Remzi H. Arpaci-Dusseau, Tyler Caraza-HarterNSDI 2025 · 11 citations
- A Dual-Agent Scheduler for Distributed Deep Learning Jobs on Public Cloud via Reinforcement LearningMingzhe Xing, Hangyu Mao, Shenglin Yin, Lichen Pan et al.KDD 2023 · 9 citations
- SMIless: Serving DAG-based Inference with Dynamic Invocations under Serverless ComputingChengzhi Lu, Huanle Xu, Yudan Li, Wenyan Chen et al.SC 2024 · 8 citations
Related papers
- Ditto: Efficient Serverless Analytics with Elastic ParallelismChao Jin, Zili Zhang, Xingyu Xiang, Songyun Zou et al.SIGCOMM 2023 · 27 citations
- ALPS: An Adaptive Learning, Priority OS Scheduler for Serverless FunctionsYuqi Fu, Ruizhe Shi, Haoliang Wang, Songqing Chen et al.USENIX ATC 2024 · 12 citations
- COSE: Configuring Serverless Functions using Statistical LearningNabeel Akhtar, Ali Raza, Vatche Ishakian, Ibrahim MattaINFOCOM 2020 · 97 citations
- Lambada: Interactive Data Analytics on Cold Data Using Serverless Cloud InfrastructureIngo Müller, Renato Marroquín, Gustavo AlonsoSIGMOD 2020 · 135 citations
- MinFlow: High-performance and Cost-efficient Data Passing for I/O-intensive Stateful Serverless AnalyticsTao Li, Yongkun Li, Wenzhe Zhu, Yinlong Xu et al.FAST 2024 · 5 citations
