NetCrafter: Tailoring Network Traffic for Non-Uniform Bandwidth Multi-GPU Systems
Amel Fatima, Yang Yang, Yifan Sun, Rachata Ausavarungnirun, Adwait Jog
Abstract
Multiple Graphics Processing Units (GPUs) are being integrated into systems to meet the computing demands of emerging workloads.To continuously support more GPUs in a system, it is important to connect them efficiently and effectively.To this end, emerging multi-GPU systems are adopting a hierarchical approach -a group of GPUs with high affinity are connected with higher-bandwidth networks, while multiple groups of GPUs are connected with lowerbandwidth networks to support the scaling of GPUs.Unfortunately, such a non-uniform bandwidth configuration leads to significant performance bottlenecks, especially across lower-bandwidth networks.We present NetCrafter, a combination of novel approaches to deal with the network traffic.NetCrafter is based on three observations: a) not all flits in the network fully utilize the network bandwidth, b) not all requested flits are even necessary -they are requested in the hope that their data might be useful later, c) some flits are more latency-sensitive than others and must be prioritized in the network.NetCrafter leverages these observations to reduce the network traffic by stitching compatible flits that are partly filled, and trimming the number of flits by not fetching flits that are unnecessary.NetCrafter also effectively manages network traffic by sequencing flits such that latency-sensitive flits reach their destinations faster.Although our proposed techniques are generic and can be applied to any network, they are especially useful in alleviating the bottlenecks presented by lower-bandwidth networks connecting multiple groups of GPUs.Overall, NetCrafter significantly improves multi-GPU performance, thereby contributing to efficient scaling of GPU-based systems.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 3ea31f23-7892-44d9-97cd-afa27fa7bff2Related papers
- NearFetch: Saving Inter-Module Bandwidth in Many-Chip-Module GPUsXia Zhao, Guangda Zhang, Lu Wang, Shiqing Zhang et al.HPCA 2025 · 3 citations
- Snake: A Variable-length Chain-based Prefetching for GPUsSaba Mostofi, Hajar Falahati, Negin Mahani, Pejman Lotfi-Kamran et al.MICRO 2023 · 11 citations
- SAC: Sharing-Aware Caching in Multi-Chip GPUsShiqing Zhang, Mahmood Naderan-Tahan, Magnus Jahre, Lieven EeckhoutISCA 2023 · 18 citations
- Memory Harvesting in Multi-GPU Systems with Hierarchical Unified Virtual MemorySangjin Choi, Taeksoo Kim, Jinwoo Jeong, Rachata Ausavarungnirun et al.USENIX ATC 2022 · 28 citations
- CHOPIN: Scalable Graphics Rendering in Multi-GPU Systems via Parallel Image CompositionXiaowei Ren, Mieszko LisHPCA 2021 · 14 citations
