ATLAS: Alibaba Dataset and Benchmark for Learning-Augmented Scheduling
Zhiyun Jiang, Tianming Zhao, Chunqiu xia, Albert Zomaya
Abstract
Learning-augmented scheduling uses ML predictions to improve decision-making under uncertainty. Many algorithms in this class have been proposed with better theoretical guarantees than the classic methods. Translating these theoretical results into practice, however, requires an understanding of real workloads. Such an understanding is hard to develop because existing production traces either lack the ground-truth processing times or are not publicly available, while synthetic benchmarks fail to represent real-world complexity. We fill this gap by introducing Alibaba Trace for Learning-Augmented Scheduling (ATLAS), a research-ready dataset derived from Alibaba's Platform of Artificial Intelligence (PAI) cluster trace—a production system that processes hundreds of thousands of ML jobs per day. The ATLAS dataset has been cleaned and features engineered to represent the inputs and constraints of non-clairvoyant scheduling, including user tags, resource requests (CPU/GPU/memory), and job structures with ground-truth processing times. We develop a prediction benchmark reporting prediction error metrics, along with feature importance analysis, and introduce a novel multiple-stage ML model. We also provide a scheduling benchmark for minimizing the total completion time, max-stretch, and makespan. ATLAS is a reproducible foundation for researchers to study learning-augmented scheduling on real workloads, available at https://github.com/zhiyunjiang0810/non-clairvoyant-with-predictions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6830f017-241d-4935-bde5-5e6d60ae0b48Builds on5
- Online Scheduling via Learned WeightsSilvio Lattanzi, Thomas Lavastida, Benjamin Moseley, Sergei VassilvitskiiSODA 2020 · 83 citations
- Shockwave: Fair and Efficient Cluster Scheduling for Dynamic Adaptation in Machine LearningPengfei Zheng, Rui Pan, Tarannum Khan, Shivaram Venkataraman et al.NSDI 2023 · 56 citations
- Non-clairvoyant Scheduling with Partial PredictionsZiyad Benomar, Vianney PerchetICML 2024 · 11 citations
- Competitive Fair Scheduling with PredictionsTianming Zhao, Chunqiu Xia, Xiaomin Chang, Chunhao Li et al.ICLR 2025
- MLaaS in the Wild: Workload Analysis and Scheduling in Large-Scale Heterogeneous GPU ClustersQizhen Weng, Wencong Xiao, Yinghao Yu, Wei Wang et al.NSDI 2022
Related papers
- Understanding and Optimizing Workloads for Unified Resource Management in Large Cloud PlatformsChengzhi Lu, Huanle Xu, Kejiang Ye, Guoyao Xu et al.EuroSys 2023 · 34 citations
- Non-Clairvoyant Scheduling with Progress BarsZiyad Benomar, Romain Cosson, Alexander Lindermayr, Jens SchlöterNeurIPS 2025 · 8 citations
- Pollux: Co-adaptive Cluster Scheduling for Goodput-Optimized Deep LearningAurick Qiao, Sang Keun Choe, Suhas Jayaram Subramanya, Willie Neiswanger et al.OSDI 2021 · 258 citations
- PBench: Workload Synthesizer with Real Statistics for Cloud Analytics BenchmarkingYan Zhou, Chunwei Liu, Bhuvan Urgaonkar, Zhengle Wang et al.VLDB 2025 · 4 citations
- Exathlon: A Benchmark for Explainable Anomaly Detection over Time SeriesVincent Jacob, Fei Song, Arnaud Stiegler, Bijan Rad et al.VLDB 2021 · 97 citations
