Scalable Irregular Parallelism with GPUs: Getting CPUs Out of the Way
Yuxin Chen, Benjamin Brock, Serban D. Porumbescu, Aydin Buluç, Katherine A. Yelick, John D. Owens
摘要
We present Atos, a dynamic scheduling framework for multi-node-GPU systems that supports PGAS-style lightweight one-sided memory operations within and between nodes. Atos's lightweight GPU-to-GPU communication enables latency hiding and can smooth the interconnection usage for bisection-limited problems. These benefits are significant for dynamic, irregular applications that often involve fine-grained communication at unpredictable times and without predetermined patterns. Some principles for high performance: ( 1) do not involve the CPU in the communication control path; (2) allow GPU communication within kernels, addressing memory consistency directly rather than relying on synchronization with the CPU; (3) perform dynamic communication aggregation when interconnections have limited bandwidth. By lowering the overhead of communication and allowing it within GPU kernels, we support large, high-utilization GPU kernels but with more frequent communication. We evaluate Atos on two irregular problems: Breadth-First-Search and PageRank. Atos outperforms the state-of-the-art graph libraries Gunrock, Groute and Galois on both single-node-multi-GPU and multi-node-GPU settings.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- OPER: Optimality-Guided Embedding Table Parallelization for Large-scale Recommendation ModelZheng Wang, Yuke Wang, Boyuan Feng, Guyue Huang 等USENIX ATC 2024 · 被引用 7 次
- KAMI: Communication-Avoiding General Matrix Multiplication within a Single GPUHemeng Wang, Yang Du, Sidu Li, Xiaowen Tian 等SC 2025 · 被引用 4 次
- WIC: Hiding Producer-Consumer Synchronization Delays with Warp-Level Interrupt-based GPU CommunicationsJiajian Zhang, Fangyu Wu, Hai Jiang, Qiufeng Wang 等USENIX ATC 2025 · 被引用 1 次
- MoE-Hub: Taming Software Complexity for Seamless MoE Overlap with Hardware-Accelerated Communication on Multi-GPU SystemsZhuoshan Zhou, Chen Zhang, Shuyi Zhang, Qijun Zhang 等ISCA 2026
相关 Paper
- FinePack: Transparently Improving the Efficiency of Fine-Grained Transfers in Multi-GPU SystemsHarini Muthukrishnan, Daniel Lustig, Oreste Villa, Thomas F. Wenisch 等HPCA 2023 · 被引用 14 次
- ScalaGraph: A Scalable Accelerator for Massively Parallel Graph ProcessingPengcheng Yao, Long Zheng, Yu Huang, Qinggang Wang 等HPCA 2022 · 被引用 30 次
- C2graph: A Compression-Collaboration Algorithm for CPU-GPU Hybrid Weighted Graph TraversalsNing Wang, Huaibei Li, Shen Su, Yu Gu 等ICDE 2026
- Atlas: Hierarchical Partitioning for Quantum Circuit Simulation on GPUsMingkuan Xu, Shiyi Cao, Xupeng Miao, Umut A. Acar 等SC 2024 · 被引用 9 次
- BLAD: Adaptive Load Balanced Scheduling and Operator Overlap Pipeline For Accelerating The Dynamic GNN TrainingKaihua Fu, Quan Chen, Yuzhuo Yang, Jiuchen Shi 等SC 2023 · 被引用 11 次
