ATLAS: Alibaba Dataset and Benchmark for Learning-Augmented Scheduling
Zhiyun Jiang, Tianming Zhao, Chunqiu xia, Albert Zomaya
摘要
Learning-augmented scheduling uses ML predictions to improve decision-making under uncertainty. Many algorithms in this class have been proposed with better theoretical guarantees than the classic methods. Translating these theoretical results into practice, however, requires an understanding of real workloads. Such an understanding is hard to develop because existing production traces either lack the ground-truth processing times or are not publicly available, while synthetic benchmarks fail to represent real-world complexity. We fill this gap by introducing Alibaba Trace for Learning-Augmented Scheduling (ATLAS), a research-ready dataset derived from Alibaba's Platform of Artificial Intelligence (PAI) cluster trace—a production system that processes hundreds of thousands of ML jobs per day. The ATLAS dataset has been cleaned and features engineered to represent the inputs and constraints of non-clairvoyant scheduling, including user tags, resource requests (CPU/GPU/memory), and job structures with ground-truth processing times. We develop a prediction benchmark reporting prediction error metrics, along with feature importance analysis, and introduce a novel multiple-stage ML model. We also provide a scheduling benchmark for minimizing the total completion time, max-stretch, and makespan. ATLAS is a reproducible foundation for researchers to study learning-augmented scheduling on real workloads, available at https://github.com/zhiyunjiang0810/non-clairvoyant-with-predictions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- Online Scheduling via Learned WeightsSilvio Lattanzi, Thomas Lavastida, Benjamin Moseley, Sergei VassilvitskiiSODA 2020 · 被引用 83 次
- Shockwave: Fair and Efficient Cluster Scheduling for Dynamic Adaptation in Machine LearningPengfei Zheng, Rui Pan, Tarannum Khan, Shivaram Venkataraman 等NSDI 2023 · 被引用 56 次
- Non-clairvoyant Scheduling with Partial PredictionsZiyad Benomar, Vianney PerchetICML 2024 · 被引用 11 次
- Competitive Fair Scheduling with PredictionsTianming Zhao, Chunqiu Xia, Xiaomin Chang, Chunhao Li 等ICLR 2025
- MLaaS in the Wild: Workload Analysis and Scheduling in Large-Scale Heterogeneous GPU ClustersQizhen Weng, Wencong Xiao, Yinghao Yu, Wei Wang 等NSDI 2022
相关 Paper
- Understanding and Optimizing Workloads for Unified Resource Management in Large Cloud PlatformsChengzhi Lu, Huanle Xu, Kejiang Ye, Guoyao Xu 等EuroSys 2023 · 被引用 34 次
- Non-Clairvoyant Scheduling with Progress BarsZiyad Benomar, Romain Cosson, Alexander Lindermayr, Jens SchlöterNeurIPS 2025 · 被引用 8 次
- Pollux: Co-adaptive Cluster Scheduling for Goodput-Optimized Deep LearningAurick Qiao, Sang Keun Choe, Suhas Jayaram Subramanya, Willie Neiswanger 等OSDI 2021 · 被引用 258 次
- PBench: Workload Synthesizer with Real Statistics for Cloud Analytics BenchmarkingYan Zhou, Chunwei Liu, Bhuvan Urgaonkar, Zhengle Wang 等VLDB 2025 · 被引用 4 次
- Exathlon: A Benchmark for Explainable Anomaly Detection over Time SeriesVincent Jacob, Fei Song, Arnaud Stiegler, Bijan Rad 等VLDB 2021 · 被引用 97 次
