Lifting and Dropping VMs to Dynamically Transition Between Time- and Space-sharing for Large-Scale HPC Systems
Nicholas Gordon, John R. Lange
Abstract
As HPC environments increasingly integrate with edge based systems, system architectures will need to handle a broader class of workloads and scheduling requirements. One result of this shift will be the need to simultaneously support bulk-synchronous parallel (BSP) and on-demand service based applications on the same infrastructure. This in turn will require that future resource management approaches utilize both space-shared as well as time-shared resource scheduling strategies. In this work we introduce the concept of "VM-lifting'' (and its inverse "VM-Dropping'') which allows dynamically switching an HPC workload between space-shared and time-shared scheduling regimes. Our work targets co-kernel based HPC system software environments, in which multiple specialized OS kernels execute natively on dedicated physical resource partitions inside a single compute node. With VM-lifting, a native co-kernel can be migrated at runtime to and from locally hosted Virtual Machine Environments due to changing scheduling requirements of the node. This allows an HPC node to be dynamically (re-)configured as either a time-shared Infrastructure-as-a-Service (IaaS) resource or a dedicated space shared resource based on the current workload demands. We have implemented this approach in the context of the Hobbes Exascale System Software stack and have demonstrated that a node can be reconfigured with minimal impact on the running applications.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 868f3b4d-9124-464e-820b-51bd167f8650Builds on1
Related papers
- HardHarvest: Hardware-Supported Core Harvesting for MicroservicesJovan Stojkovic, Chunao Liu, Muhammad Shahbaz, Josep TorrellasISCA 2025 · 4 citations
- Optimizing Task Scheduling in Cloud VMs with Accurate vCPU AbstractionEdward Guo, Weiwei Jia, Xiaoning Ding, Jianchen ShanEuroSys 2025 · 3 citations
- Memory-harvesting VMs in cloud platformsAlexander Fuerst, Stanko Novakovic, Iñigo Goiri, Gohar Irfan Chaudhry et al.ASPLOS 2022 · 39 citations
- Using Unused: Non-Invasive Dynamic FaaS Infrastructure with HPC-WhiskBartlomiej Przybylski, Maciej Pawlik, Pawel Zuk, Bartlomiej Lagosz et al.SC 2022 · 9 citations
- Fast Core Scheduling with Userspace Process AbstractionJiazhen Lin, Youmin Chen, Shiwei Gao, Youyou LuSOSP 2024 · 3 citations
