GICC: A High-Performance Runtime for GPU-Initiated Communication and Coordination in Modern HPC Systems
Baodi Shan, Mauricio Araya-Polo, Barbara M. Chapman
Abstract
An increasingly popular approach for distributed GPU applications is kernel-level, cross-node coordination to reduce launch overheads and improve compute-communication overlap. The reality is that such support is lacking. On one hand, on OFI-based interconnects such as HPE Slingshot-which powers six of the top ten systems in the November 2025 Top500, including the top three-GPU kernels cannot autonomously drive distributed coordination: existing runtimes rely on host-driven progress and do not provide a bounded mechanism for recycling pre-staged NIC work across repeated GPU-triggered operations. On the other hand, on InfiniBand, GPUinitiated communication is possible, but current implementations incur unnecessary synchronization and locking overheads. This paper presents GICC, a GPU-driven coordination framework that enables GPU kernels to issue coordination commands and directly trigger NIC-level operations without host involvement on the fast path. For example, in stencil computations, GPU threads can directly initiate halo exchanges with neighboring nodes as soon as boundary regions are computed, rather than waiting for the kernel to complete and relying on the host to orchestrate communication. This eliminates synchronization overhead and enables fine-grained overlap between interior computation and boundary data transfer. GICC decouples coordination semantics from data movement and introduces an asynchronous resource reclamation scheme: the NIC signals completion to both GPU and host memory, enabling a lightweight host thread to recycle NIC resources concurrently with GPU execution-without injecting latency into the coordination path. This design enables sustained GPU-driven coordination under finite NIC state, a capability absent from existing runtimes on OFI-based fabrics.
We implement GICC on NVIDIA and AMD GPUs over both InfiniBand and Slingshot (OFI). On Slingshot, GICC reduces average per-coordination latency by up to 229× and improves weak scaling efficiency by up to 25%. On InfiniBand, GICC achieves up to 1.95× lower put latency than NVSHMEM by eliminating unnecessary locking and synchronization. For an industrial stencil-based proxy application on 64 AMD MI250X GCDs, GPU-aware MPI incurs over 52% higher communication time than GICC, and GICC achieves 42% parallel efficiency versus MPI's 35.4%. These results demonstrate that GPU-driven, resource-aware coordination is both necessary This work is licensed under a Creative Commons Attribution 4.0 International License. HPDC '26, Cleveland, OH, USA
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ec358ed9-299a-4493-b9da-97dc58247cd8Builds on1
Related papers
- WIC: Hiding Producer-Consumer Synchronization Delays with Warp-Level Interrupt-based GPU CommunicationsJiajian Zhang, Fangyu Wu, Hai Jiang, Qiufeng Wang et al.USENIX ATC 2025 · 1 citation
- MGI: A Communication Framework for Data Processing in Massive GPU InfrastructuresDi Wu, Hongshi Tan, Hanzhang Yang, Bingsheng He et al.VLDB 2026
- <u>G</u>PU <u>i</u>nitiated <u>O</u>penSHMEM: correct and efficient intra-kernel networking for dGPUsKhaled Hamidouche, Michael LeBeanePPoPP 2020 · 17 citations
- Asynchrony and GPUs: Bridging this Dichotomy for I/O with AGIOJihoon Han, Anand Sivasubramaniam, Chia-Hao Chang, Vikram Sharma Mailthody et al.ASPLOS 2026 · 1 citation
- Scaling GPU-to-CPU Migration for Efficient Distributed Execution on CPU ClustersRuobing Han, Hyesoon KimPPoPP 2026
