USENIX ATC2025顶会
WIC: Hiding Producer-Consumer Synchronization Delays with Warp-Level Interrupt-based GPU Communications
Jiajian Zhang, Fangyu Wu, Hai Jiang, Qiufeng Wang, Genlang Chen, Chaoyi Pang
摘要
GPU communication plays a pivotal role in collaborative computation across multiple devices. Despite advancements in inter-device communication fabrics and architectures, synchronization still remains a significant challenge due to the manual coordination required between producers and consumers at the application level. In this work, we first reveal that traditional synchronization is a primary bottleneck in GPU communication, where consumers frequently poll for producer data availability. Specifically, early-started polling leads to the unnecessary occupation of computational resources. To address this issue, we propose Warplevel Interrupt-based Communication (WIC), a novel synchronization framework for GPU communication that introduces a fine-grained interruption mechanism at the warp level to replace repetitive polling. WIC preemptively stalls warps engaged in frequent polling and releases computational resources for other warps, thereby effectively overlapping producer-consumer synchronization with ongoing computations. Comprehensive experiments demonstrate that WIC significantly outperforms conventional polling methods by 1.13× on average across various applications with diverse communication patterns.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper16
- Griffin: Hardware-Software Support for Efficient Page Migration in Multi-GPU SystemsTrinayan Baruah, Yifan Sun, Ali Tolga Dinçer, Saiful A. Mojumder 等HPCA 2020 · 被引用 50 次
- MGG: Accelerating Graph Neural Networks with Fine-Grained Intra-Kernel Communication-Computation Pipelining on Multi-GPU PlatformsYuke Wang, Boyuan Feng, Zheng Wang, Tong Geng 等OSDI 2023 · 被引用 46 次
- HMG: Extending Cache Coherence Protocols Across Modern Hierarchical Multi-GPU SystemsXiaowei Ren, Daniel Lustig, Evgeny Bolotin, Aamer Jaleel 等HPCA 2020 · 被引用 38 次
- Accelerating the Training of Large Language Models using Efficient Activation Rematerialization and Optimal Hybrid ParallelismTailing Yuan, Yuliang Liu, Xucheng Ye, Shenglong Zhang 等USENIX ATC 2024 · 被引用 34 次
- Fast and Efficient Model Serving Using Multi-GPUs with Direct-Host-AccessJinwoo Jeong, Seungsu Baek, Jeongseob AhnEuroSys 2023 · 被引用 28 次
相关 Paper
- GICC: A High-Performance Runtime for GPU-Initiated Communication and Coordination in Modern HPC SystemsBaodi Shan, Mauricio Araya-Polo, Barbara M. ChapmanHPDC 2026
- MGI: A Communication Framework for Data Processing in Massive GPU InfrastructuresDi Wu, Hongshi Tan, Hanzhang Yang, Bingsheng He 等VLDB 2026
- A Simple Cache Coherence Scheme for Integrated CPU-GPU SystemsArdhi Wiratama Baskara Yudha, Reza Pulungan, Henry Hoffmann, Yan SolihinDAC 2020 · 被引用 3 次
- Tally: Non-Intrusive Performance Isolation for Concurrent Deep Learning WorkloadsWei Zhao, Anand Jayarajan, Gennady PekhimenkoASPLOS 2025 · 被引用 4 次
- WASP: Exploiting GPU Pipeline Parallelism with Hardware-Accelerated Automatic Warp SpecializationNeal Clayton Crago, Sana Damani, Karthikeyan Sankaralingam, Stephen W. KecklerHPCA 2024 · 被引用 13 次
