Towards Scalable Resource Management for Supercomputers
Yiqin Dai, Yong Dong, Kai Lu, Ruibo Wang, Wei Zhang, Juan Chen, Mingtian Shao, Zheng Wang
摘要
Today's supercomputers offer massive computation resources to execute a large number of user jobs. Effectively managing such large-scale hardware parallelism and workloads is essential for supercomputers. However, existing HPC resource management (RM) systems fail to capitalize on the hardware parallelism by following a centralized design used decades ago. They give poor scalability and inefficient performance on today's supercomputers, which will worsen in exascale computing. We present ESlurm, a better RM for supercomputers. As a departure from existing HPC RMs, ESlurm implements a distributed communication structure. It employs a new communication tree strategy and uses job runtime estimation to improve communications and job scheduling efficiency. ESlurm is deployed into production in a real supercomputer. We evaluate ESlurm on up to 20K nodes. Compared to state-of-the-art RM solutions, ESlurm exhibits better scalability, significantly reducing the resource usage of master nodes and improving data transfer and job scheduling efficiency by a large margin.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper2
- Enabling and scaling the HPCG benchmark on the newest generation Sunway supercomputer with 42 million heterogeneous coresQianchao Zhu, Hao Luo, Chao Yang, Mingshuo Ding 等SC 2021 · 被引用 44 次
- Revealing power, energy and thermal dynamics of a 200PF pre-exascale supercomputerWoong Shin, Vladyslav Oles, Ahmad Maroof Karimi, J. Austin Ellis 等SC 2021 · 被引用 41 次
相关 Paper
- Iris: allocation banking and identity and access management for the exascale eraGabor Torok, Mark R. Day, Rebecca Hartman-Baker, Cory SnavelySC 2020 · 被引用 4 次
- Efficient all-to-all Collective Communication Schedules for Direct-connect TopologiesPrithwish Basu, Liangyu Zhao, Jason Fantl, Siddharth Pal 等HPDC 2024 · 被引用 7 次
- Communication Algorithm-Architecture Co-Design for Distributed Deep LearningJiayi Huang, Pritam Majumder, Sungkeun Kim, Abdullah Muzahid 等ISCA 2021 · 被引用 44 次
- Concealing Compression-accelerated I/O for HPC Applications through In Situ Task SchedulingSian Jin, Sheng Di, Frédéric Vivien, Daoce Wang 等EuroSys 2024 · 被引用 13 次
- Efficient Pre-Training of LLMs via Topology-Aware Communication Alignment on More Than 9600 GPUsGuoliang He, Youhe Jiang, Wencong Xiao, Kaihua Jiang 等NeurIPS 2025 · 被引用 10 次
