GICC: A High-Performance Runtime for GPU-Initiated Communication and Coordination in Modern HPC Systems
Baodi Shan, Mauricio Araya-Polo, Barbara M. Chapman
摘要
An increasingly popular approach for distributed GPU applications is kernel-level, cross-node coordination to reduce launch overheads and improve compute-communication overlap. The reality is that such support is lacking. On one hand, on OFI-based interconnects such as HPE Slingshot-which powers six of the top ten systems in the November 2025 Top500, including the top three-GPU kernels cannot autonomously drive distributed coordination: existing runtimes rely on host-driven progress and do not provide a bounded mechanism for recycling pre-staged NIC work across repeated GPU-triggered operations. On the other hand, on InfiniBand, GPUinitiated communication is possible, but current implementations incur unnecessary synchronization and locking overheads. This paper presents GICC, a GPU-driven coordination framework that enables GPU kernels to issue coordination commands and directly trigger NIC-level operations without host involvement on the fast path. For example, in stencil computations, GPU threads can directly initiate halo exchanges with neighboring nodes as soon as boundary regions are computed, rather than waiting for the kernel to complete and relying on the host to orchestrate communication. This eliminates synchronization overhead and enables fine-grained overlap between interior computation and boundary data transfer. GICC decouples coordination semantics from data movement and introduces an asynchronous resource reclamation scheme: the NIC signals completion to both GPU and host memory, enabling a lightweight host thread to recycle NIC resources concurrently with GPU execution-without injecting latency into the coordination path. This design enables sustained GPU-driven coordination under finite NIC state, a capability absent from existing runtimes on OFI-based fabrics.
We implement GICC on NVIDIA and AMD GPUs over both InfiniBand and Slingshot (OFI). On Slingshot, GICC reduces average per-coordination latency by up to 229× and improves weak scaling efficiency by up to 25%. On InfiniBand, GICC achieves up to 1.95× lower put latency than NVSHMEM by eliminating unnecessary locking and synchronization. For an industrial stencil-based proxy application on 64 AMD MI250X GCDs, GPU-aware MPI incurs over 52% higher communication time than GICC, and GICC achieves 42% parallel efficiency versus MPI's 35.4%. These results demonstrate that GPU-driven, resource-aware coordination is both necessary This work is licensed under a Creative Commons Attribution 4.0 International License. HPDC '26, Cleveland, OH, USA
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper1
相关 Paper
- WIC: Hiding Producer-Consumer Synchronization Delays with Warp-Level Interrupt-based GPU CommunicationsJiajian Zhang, Fangyu Wu, Hai Jiang, Qiufeng Wang 等USENIX ATC 2025 · 被引用 1 次
- MGI: A Communication Framework for Data Processing in Massive GPU InfrastructuresDi Wu, Hongshi Tan, Hanzhang Yang, Bingsheng He 等VLDB 2026
- <u>G</u>PU <u>i</u>nitiated <u>O</u>penSHMEM: correct and efficient intra-kernel networking for dGPUsKhaled Hamidouche, Michael LeBeanePPoPP 2020 · 被引用 17 次
- Asynchrony and GPUs: Bridging this Dichotomy for I/O with AGIOJihoon Han, Anand Sivasubramaniam, Chia-Hao Chang, Vikram Sharma Mailthody 等ASPLOS 2026 · 被引用 1 次
- Scaling GPU-to-CPU Migration for Efficient Distributed Execution on CPU ClustersRuobing Han, Hyesoon KimPPoPP 2026
