SC2022Top-tier venue
Towards Scalable Resource Management for Supercomputers
Yiqin Dai, Yong Dong, Kai Lu, Ruibo Wang, Wei Zhang, Juan Chen, Mingtian Shao, Zheng Wang
Abstract
Today's supercomputers offer massive computation resources to execute a large number of user jobs. Effectively managing such large-scale hardware parallelism and workloads is essential for supercomputers. However, existing HPC resource management (RM) systems fail to capitalize on the hardware parallelism by following a centralized design used decades ago. They give poor scalability and inefficient performance on today's supercomputers, which will worsen in exascale computing. We present ESlurm, a better RM for supercomputers. As a departure from existing HPC RMs, ESlurm implements a distributed communication structure. It employs a new communication tree strategy and uses job runtime estimation to improve communications and job scheduling efficiency. ESlurm is deployed into production in a real supercomputer. We evaluate ESlurm on up to 20K nodes. Compared to state-of-the-art RM solutions, ESlurm exhibits better scalability, significantly reducing the resource usage of master nodes and improving data transfer and job scheduling efficiency by a large margin.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d04f5ae6-4ae2-4ba3-987e-d1eb187d9c7cCited by top-tier papers1
Ask how each one uses itBuilds on2
- Enabling and scaling the HPCG benchmark on the newest generation Sunway supercomputer with 42 million heterogeneous coresQianchao Zhu, Hao Luo, Chao Yang, Mingshuo Ding et al.SC 2021 · 44 citations
- Revealing power, energy and thermal dynamics of a 200PF pre-exascale supercomputerWoong Shin, Vladyslav Oles, Ahmad Maroof Karimi, J. Austin Ellis et al.SC 2021 · 41 citations
Related papers
- Iris: allocation banking and identity and access management for the exascale eraGabor Torok, Mark R. Day, Rebecca Hartman-Baker, Cory SnavelySC 2020 · 4 citations
- Efficient all-to-all Collective Communication Schedules for Direct-connect TopologiesPrithwish Basu, Liangyu Zhao, Jason Fantl, Siddharth Pal et al.HPDC 2024 · 7 citations
- Communication Algorithm-Architecture Co-Design for Distributed Deep LearningJiayi Huang, Pritam Majumder, Sungkeun Kim, Abdullah Muzahid et al.ISCA 2021 · 44 citations
- Concealing Compression-accelerated I/O for HPC Applications through In Situ Task SchedulingSian Jin, Sheng Di, Frédéric Vivien, Daoce Wang et al.EuroSys 2024 · 13 citations
- Efficient Pre-Training of LLMs via Topology-Aware Communication Alignment on More Than 9600 GPUsGuoliang He, Youhe Jiang, Wencong Xiao, Kaihua Jiang et al.NeurIPS 2025 · 10 citations
