SC2022Top-tier venue
Scalable Irregular Parallelism with GPUs: Getting CPUs Out of the Way
Yuxin Chen, Benjamin Brock, Serban D. Porumbescu, Aydin Buluç, Katherine A. Yelick, John D. Owens
Abstract
We present Atos, a dynamic scheduling framework for multi-node-GPU systems that supports PGAS-style lightweight one-sided memory operations within and between nodes. Atos's lightweight GPU-to-GPU communication enables latency hiding and can smooth the interconnection usage for bisection-limited problems. These benefits are significant for dynamic, irregular applications that often involve fine-grained communication at unpredictable times and without predetermined patterns. Some principles for high performance: ( 1) do not involve the CPU in the communication control path; (2) allow GPU communication within kernels, addressing memory consistency directly rather than relying on synchronization with the CPU; (3) perform dynamic communication aggregation when interconnections have limited bandwidth. By lowering the overhead of communication and allowing it within GPU kernels, we support large, high-utilization GPU kernels but with more frequent communication. We evaluate Atos on two irregular problems: Breadth-First-Search and PageRank. Atos outperforms the state-of-the-art graph libraries Gunrock, Groute and Galois on both single-node-multi-GPU and multi-node-GPU settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- OPER: Optimality-Guided Embedding Table Parallelization for Large-scale Recommendation ModelZheng Wang, Yuke Wang, Boyuan Feng, Guyue Huang et al.USENIX ATC 2024 · 7 citations
- KAMI: Communication-Avoiding General Matrix Multiplication within a Single GPUHemeng Wang, Yang Du, Sidu Li, Xiaowen Tian et al.SC 2025 · 4 citations
- WIC: Hiding Producer-Consumer Synchronization Delays with Warp-Level Interrupt-based GPU CommunicationsJiajian Zhang, Fangyu Wu, Hai Jiang, Qiufeng Wang et al.USENIX ATC 2025 · 1 citation
- MoE-Hub: Taming Software Complexity for Seamless MoE Overlap with Hardware-Accelerated Communication on Multi-GPU SystemsZhuoshan Zhou, Chen Zhang, Shuyi Zhang, Qijun Zhang et al.ISCA 2026
Related papers
- FinePack: Transparently Improving the Efficiency of Fine-Grained Transfers in Multi-GPU SystemsHarini Muthukrishnan, Daniel Lustig, Oreste Villa, Thomas F. Wenisch et al.HPCA 2023 · 14 citations
- ScalaGraph: A Scalable Accelerator for Massively Parallel Graph ProcessingPengcheng Yao, Long Zheng, Yu Huang, Qinggang Wang et al.HPCA 2022 · 30 citations
- C2graph: A Compression-Collaboration Algorithm for CPU-GPU Hybrid Weighted Graph TraversalsNing Wang, Huaibei Li, Shen Su, Yu Gu et al.ICDE 2026
- Atlas: Hierarchical Partitioning for Quantum Circuit Simulation on GPUsMingkuan Xu, Shiyi Cao, Xupeng Miao, Umut A. Acar et al.SC 2024 · 9 citations
- BLAD: Adaptive Load Balanced Scheduling and Operator Overlap Pipeline For Accelerating The Dynamic GNN TrainingKaihua Fu, Quan Chen, Yuzhuo Yang, Jiuchen Shi et al.SC 2023 · 11 citations
