Scaling GPU-to-CPU Migration for Efficient Distributed Execution on CPU Clusters
Ruobing Han, Hyesoon Kim
Abstract
The growing demand for GPU resources has led to widespread shortages in data centers, prompting the exploration of CPUs as an alternative for executing GPU programs. While prior research supports executing GPU programs on single CPUs, these approaches struggle to achieve competitive performance due to the computational capacity gap between GPUs and CPUs.
To further improve performance, we introduce CuCC, a framework that scales GPU-to-CPU migration to CPU clusters and utilizes distributed CPU nodes to execute GPU programs. Compared to single-CPU execution, CPU cluster execution requires cross-node communication to maintain data consistency. We present the CuCC execution workflow and communication optimizations, which aim to reduce network overhead. Evaluations demonstrate that CuCC achieves high scalability on large-scale CPU clusters and delivers runtimes approaching those of GPUs. In terms of cluster-wide throughput, CuCC enables CPUs to achieve an average of 2.59× higher throughput than GPUs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bd5f0457-9376-4f2c-a0e9-43fb43d909b1Builds on4
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Performance GPU-to-CPU Transpilation and Optimization via High-Level Parallel ConstructsWilliam S. Moses, Ivan R. Ivanov, Jens Domke, Toshio Endo et al.PPoPP 2023 · 27 citations
- Dopia: online parallelism management for integrated CPU/GPU architecturesYounghyun Cho, Jiyeon Park, Florian Negele, Changyeon Jo et al.PPoPP 2022 · 9 citations
- Unleashing CPU Potential for Executing GPU Programs Through Compiler/Runtime OptimizationsRuobing Han, Jisheng Zhao, Hyesoon KimMICRO 2024 · 4 citations
Related papers
- MGI: A Communication Framework for Data Processing in Massive GPU InfrastructuresDi Wu, Hongshi Tan, Hanzhang Yang, Bingsheng He et al.VLDB 2026
- Griffin: Hardware-Software Support for Efficient Page Migration in Multi-GPU SystemsTrinayan Baruah, Yifan Sun, Ali Tolga Dinçer, Saiful A. Mojumder et al.HPCA 2020 · 50 citations
- GICC: A High-Performance Runtime for GPU-Initiated Communication and Coordination in Modern HPC SystemsBaodi Shan, Mauricio Araya-Polo, Barbara M. ChapmanHPDC 2026
- HARP: Orchestrating Automated Parallel Training on Heterogeneous GPU ClustersAntian Liang, Zhigang Zhao, Kai Zhang, Xuri Shi et al.EuroSys 2026 · 1 citation
- CPU- and GPU-initiated Communication Strategies for Conjugate Gradient Methods on Large GPU ClustersJames D. Trotter, Sinan Ekmekçibasi, Dogan Sagbili, Johannes Langguth et al.SC 2025 · 2 citations
