SPADE: Signal-Aware DAG Scheduling and Dynamic Provisioning for Data Processing Clusters
Adam Lechowicz, Rohan Shenoy, Noman Bashir, Mohammad Hajiesmaili, Adam Wierman, Christina Delimitrou
Abstract
As AI-driven demand reshapes the data center landscape, external signals—such as energy cost, carbon intensity, power availability, and water usage—are increasingly dictating how much compute is available at any moment. These signals tend to vary over time, challenging traditional cluster schedulers, which implicitly assume stable resource supply, and calls for systems that continuously adapt to time-varying conditions. We focus on batch data-processing workloads, which are delay-tolerant but constitute a healthy fraction of total compute, making them a natural target for such flexibility. The directed acyclic graph (DAG) structure of these data-processing jobs makes decisions uniquely challenging, since delaying certain tasks in the DAG (e.g., bottleneck tasks) can stall entire pipelines. We introduce SPADE , a signal-aware scheduling and provisioning system that jointly considers workload DAG structure and external time-varying signals when deciding how (provisioning) and when (scheduling) to allocate resources. To underscore the importance of coupling these decisions, we evaluate SAP , an ablated system that preserves SPADE ’s signal-aware provisioning but delegates scheduling to arbitrary signal-agnostic policies. Using a Spark prototype deployed on a 100-node Kubernetes cluster, we show that SPADE reduces a secondary objective (e.g., the cost associated with carbon intensity or energy price) by 32.9% while maintaining overall cluster throughput.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e54ad719-28ef-48e4-8915-eb2e14beb8b5Builds on9
- Lessons Learned from the Chameleon TestbedKate Keahey, Jason Anderson, Zhuo Zhen, Pierre Riteau et al.USENIX ATC 2020 · 398 citations
- Carbon Explorer: A Holistic Framework for Designing Carbon Aware DatacentersBilge Acun, Benjamin C. Lee, Fiodar Kazhamiaka, Kiwan Maeng et al.ASPLOS 2023 · 171 citations
- Data Center Power Oversubscription with a Medium Voltage Power Plane and Priority-Aware CappingVarun Sakalkar, Vasileios Kontorinis, David Landhuis, Shaohong Li et al.ASPLOS 2020 · 56 citations
- Going Green for Less Green: Optimizing the Cost of Reducing Cloud Carbon EmissionsWalid A. Hanafy, Qianlin Liang, Noman Bashir, Abel Souza et al.ASPLOS 2024 · 43 citations
- Minimalistic Predictions to Schedule Jobs with Online Precedence ConstraintsAlexandra Anna Lassota, Alexander Lindermayr, Nicole Megow, Jens SchlöterICML 2023 · 17 citations
Related papers
- WaterWise: Co-optimizing Carbon- and Water-Footprint Toward Environmentally Sustainable Cloud ComputingYankai Jiang, Rohan Basu Roy, Raghavendra Kanakagiri, Devesh TiwariPPoPP 2025 · 17 citations
- On the Limitations of Carbon-Aware Temporal and Spatial Workload Shifting in the CloudThanathorn Sukprasert, Abel Souza, Noman Bashir, David Irwin et al.EuroSys 2024 · 73 citations
- A House United Within Itself: SLO-Awareness for On-Premises Containerized ML Inference Clusters via FaroBeomyeol Jeon, Chen Wang, Diana Arroyo, Alaa Youssef et al.EuroSys 2025
- GREEN: Carbon-efficient Resource Scheduling for Machine Learning ClustersKaiqiang Xu, Decang Sun, Han Tian, Junxue Zhang et al.NSDI 2025 · 23 citations
- Tailored Learning-Based Scheduling for Kubernetes-Oriented Edge-Cloud SystemYiwen Han, Shihao Shen, Xiaofei Wang, Shiqiang Wang et al.INFOCOM 2021 · 93 citations
