Lune

ISCA2026Top-tier venue

RoCC: Harnessing Raster Operations Pipeline for Efficient Tensor Collective Communication

Yuan Feng, Daniel Wong, Hyeran Jeon

2026Year

Abstract

This paper introduces RoCC, which enables finegrained overlapping between compute and collective communication (CC) phases of LLM computing on GPUs, by offloading the CC to underutilized raster operations pipelines (ROPs). ROPs can provide fruitful performance for CC as they reside near the memory and have reduction computation capability. We first reverse engineer the ROP microarchitecture of two GPU architectures to model ROPs and add small logics to enable asynchronous computing and messaging for CC. We also decompose any CC operations into a sequence of ROP microoperations. In our cycle-level simulations of a 4- to 8-GPU node with LLM training workloads, RoCC delivers an average of 51% and 23% speedups over the non-overlapping baseline and oracle kernel fusion, with only 2.4%L2\mathbf{2. 4 \%} \mathbf{L2} cache worth of area. On larger systems with 32 to 256 GPUs, RoCC consistently achieves speedups from 13-21%.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get be5cfca9-4dd6-474e-8b72-3b5f273847ef

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines