RoCC: Harnessing Raster Operations Pipeline for Efficient Tensor Collective Communication
Yuan Feng, Daniel Wong, Hyeran Jeon
摘要
This paper introduces RoCC, which enables finegrained overlapping between compute and collective communication (CC) phases of LLM computing on GPUs, by offloading the CC to underutilized raster operations pipelines (ROPs). ROPs can provide fruitful performance for CC as they reside near the memory and have reduction computation capability. We first reverse engineer the ROP microarchitecture of two GPU architectures to model ROPs and add small logics to enable asynchronous computing and messaging for CC. We also decompose any CC operations into a sequence of ROP microoperations. In our cycle-level simulations of a 4- to 8-GPU node with LLM training workloads, RoCC delivers an average of 51% and 23% speedups over the non-overlapping baseline and oracle kernel fusion, with only cache worth of area. On larger systems with 32 to 256 GPUs, RoCC consistently achieves speedups from 13-21%.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Multipath Collective Communication Beyond Scale-up Networks in GPU CloudsYuchen Xu, Jianglong Nie, Baojia Li, Mingzhuo Chen 等EuroSys 2026 · 被引用 1 次
- Optimizing Distributed ML Communication with Fused Computation-Collective OperationsKishore Punniyamurthy, Khaled Hamidouche, Bradford M. BeckmannSC 2024 · 被引用 11 次
- Towards Compute-Aware In-Switch Computing for LLMs Tensor-Parallelism on Multi-GPU SystemsChen Zhang, Qijun Zhang, Zhuoshan Zhou, Yijia Diao 等HPCA 2026 · 被引用 1 次
- Centauri: Enabling Efficient Scheduling for Communication-Computation Overlap in Large Model Training via Communication PartitioningChang Chen, Xiuhong Li, Qianchao Zhu, Jiangfei Duan 等ASPLOS 2024 · 被引用 52 次
- Compass: Dissecting Communication and Computation Operators for Efficient LLM TrainingGuangyu Xiang, Lin Zhang, Haoxuan Yu, Xinglin Pan 等INFOCOM 2026
