Lune

HPDC2026顶会

GICC: A High-Performance Runtime for GPU-Initiated Communication and Coordination in Modern HPC Systems

Baodi Shan, Mauricio Araya-Polo, Barbara M. Chapman

2026年份

摘要

An increasingly popular approach for distributed GPU applications is kernel-level, cross-node coordination to reduce launch overheads and improve compute-communication overlap. The reality is that such support is lacking. On one hand, on OFI-based interconnects such as HPE Slingshot-which powers six of the top ten systems in the November 2025 Top500, including the top three-GPU kernels cannot autonomously drive distributed coordination: existing runtimes rely on host-driven progress and do not provide a bounded mechanism for recycling pre-staged NIC work across repeated GPU-triggered operations. On the other hand, on InfiniBand, GPUinitiated communication is possible, but current implementations incur unnecessary synchronization and locking overheads. This paper presents GICC, a GPU-driven coordination framework that enables GPU kernels to issue coordination commands and directly trigger NIC-level operations without host involvement on the fast path. For example, in stencil computations, GPU threads can directly initiate halo exchanges with neighboring nodes as soon as boundary regions are computed, rather than waiting for the kernel to complete and relying on the host to orchestrate communication. This eliminates synchronization overhead and enables fine-grained overlap between interior computation and boundary data transfer. GICC decouples coordination semantics from data movement and introduces an asynchronous resource reclamation scheme: the NIC signals completion to both GPU and host memory, enabling a lightweight host thread to recycle NIC resources concurrently with GPU execution-without injecting latency into the coordination path. This design enables sustained GPU-driven coordination under finite NIC state, a capability absent from existing runtimes on OFI-based fabrics.

We implement GICC on NVIDIA and AMD GPUs over both InfiniBand and Slingshot (OFI). On Slingshot, GICC reduces average per-coordination latency by up to 229× and improves weak scaling efficiency by up to 25%. On InfiniBand, GICC achieves up to 1.95× lower put latency than NVSHMEM by eliminating unnecessary locking and synchronization. For an industrial stencil-based proxy application on 64 AMD MI250X GCDs, GPU-aware MPI incurs over 52% higher communication time than GICC, and GICC achieves 42% parallel efficiency versus MPI's 35.4%. These results demonstrate that GPU-driven, resource-aware coordination is both necessary This work is licensed under a Creative Commons Attribution 4.0 International License. HPDC '26, Cleveland, OH, USA

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper1

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖