Efficient and Adaptable Overlapping for Computation and Communication via Signaling and Reordering
Ke Hong, Xiuhong Li, Minxu Liu, Qiuli Mao, Tianqi Wu, Zixiao Huang, Lufang Chen, Zhong Wang, Yichong Zhang, Zhenhua Zhu, Guohao Dai, Yu Wang
摘要
Generative models have achieved remarkable success across various applications, driving the demand for multi-GPU computing. Inter-GPU communication becomes a bottleneck in multi-GPU computing systems, particularly on consumergrade GPUs. By exploiting concurrent hardware execution, overlapping computation and communication latency becomes an effective technique for mitigating the communication overhead. We identify that an efficient and adaptable overlapping design should satisfy (1) tile-wise overlapping to maximize the overlapping opportunity, (2) interference-free computation to maintain the original computational performance, and (3) communication agnosticism to reduce the development burden against varying communication primitives. Nevertheless, current designs fail to simultaneously optimize for all of those features.
To address the issue, we propose an overlapping design, named FlashOverlap, characterized by tile-wise overlapping, interference-free computation, and communication agnosticism. FlashOverlap utilizes a novel signaling mechanism: when part of the output finishes, the computation kernel sends a signal to trigger the communication of that part, while continuing the computation of the remaining part (interference-free computation). Consequently, the communication of the finished part and the computation of *Corresponding to Yu Wang
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Syncopate: Efficient Multi-GPU AI Kernels via Automatic Chunk-Centric Compute-Communication OverlapXinwei Qiang, Yue Guan, Zhengding Hu, Keren Zhou 等OSDI 2026 · 被引用 3 次
- Disdp: Disaggregating Compute, Network, and Storage for Model-Sharded Data-Parallel TrainingMo Sun, Zihan Yang, Changyue Liao, Yingtao Li 等ISCA 2026 · 被引用 1 次
- MoE-Hub: Taming Software Complexity for Seamless MoE Overlap with Hardware-Accelerated Communication on Multi-GPU SystemsZhuoshan Zhou, Chen Zhang, Shuyi Zhang, Qijun Zhang 等ISCA 2026
它引用的顶会 Paper15
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen 等ICLR 2021 · 被引用 1,954 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- Make-A-Video: Text-to-Video Generation without Text-Video DataUriel Singer, Adam Polyak, Thomas Hayes, Xi Yin 等ICLR 2023 · 被引用 313 次
- CogVideo: Large-scale Pretraining for Text-to-Video Generation via TransformersWenyi Hong, Ming Ding, Wendi Zheng, Xinghan Liu 等ICLR 2023 · 被引用 116 次
- FasterMoE: modeling and optimizing training of large-scale dynamic pre-trained modelsJiaao He, Jidong Zhai, Tiago Antunes, Haojie Wang 等PPoPP 2022 · 被引用 97 次
相关 Paper
- T3: Transparent Tracking & Triggering for Fine-grained Overlap of Compute & CollectivesSuchita Pati, Shaizeen Aga, Mahzabeen Islam, Nuwan Jayasena 等ASPLOS 2024 · 被引用 13 次
- Optimizing Distributed ML Communication with Fused Computation-Collective OperationsKishore Punniyamurthy, Khaled Hamidouche, Bradford M. BeckmannSC 2024 · 被引用 11 次
- MGG: Accelerating Graph Neural Networks with Fine-Grained Intra-Kernel Communication-Computation Pipelining on Multi-GPU PlatformsYuke Wang, Boyuan Feng, Zheng Wang, Tong Geng 等OSDI 2023 · 被引用 46 次
- WIC: Hiding Producer-Consumer Synchronization Delays with Warp-Level Interrupt-based GPU CommunicationsJiajian Zhang, Fangyu Wu, Hai Jiang, Qiufeng Wang 等USENIX ATC 2025 · 被引用 1 次
- RoCC: Harnessing Raster Operations Pipeline for Efficient Tensor Collective CommunicationYuan Feng, Daniel Wong, Hyeran JeonISCA 2026
