NetCrafter: Tailoring Network Traffic for Non-Uniform Bandwidth Multi-GPU Systems
Amel Fatima, Yang Yang, Yifan Sun, Rachata Ausavarungnirun, Adwait Jog
摘要
Multiple Graphics Processing Units (GPUs) are being integrated into systems to meet the computing demands of emerging workloads.To continuously support more GPUs in a system, it is important to connect them efficiently and effectively.To this end, emerging multi-GPU systems are adopting a hierarchical approach -a group of GPUs with high affinity are connected with higher-bandwidth networks, while multiple groups of GPUs are connected with lowerbandwidth networks to support the scaling of GPUs.Unfortunately, such a non-uniform bandwidth configuration leads to significant performance bottlenecks, especially across lower-bandwidth networks.We present NetCrafter, a combination of novel approaches to deal with the network traffic.NetCrafter is based on three observations: a) not all flits in the network fully utilize the network bandwidth, b) not all requested flits are even necessary -they are requested in the hope that their data might be useful later, c) some flits are more latency-sensitive than others and must be prioritized in the network.NetCrafter leverages these observations to reduce the network traffic by stitching compatible flits that are partly filled, and trimming the number of flits by not fetching flits that are unnecessary.NetCrafter also effectively manages network traffic by sequencing flits such that latency-sensitive flits reach their destinations faster.Although our proposed techniques are generic and can be applied to any network, they are especially useful in alleviating the bottlenecks presented by lower-bandwidth networks connecting multiple groups of GPUs.Overall, NetCrafter significantly improves multi-GPU performance, thereby contributing to efficient scaling of GPU-based systems.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- NearFetch: Saving Inter-Module Bandwidth in Many-Chip-Module GPUsXia Zhao, Guangda Zhang, Lu Wang, Shiqing Zhang 等HPCA 2025 · 被引用 3 次
- Snake: A Variable-length Chain-based Prefetching for GPUsSaba Mostofi, Hajar Falahati, Negin Mahani, Pejman Lotfi-Kamran 等MICRO 2023 · 被引用 11 次
- SAC: Sharing-Aware Caching in Multi-Chip GPUsShiqing Zhang, Mahmood Naderan-Tahan, Magnus Jahre, Lieven EeckhoutISCA 2023 · 被引用 18 次
- Memory Harvesting in Multi-GPU Systems with Hierarchical Unified Virtual MemorySangjin Choi, Taeksoo Kim, Jinwoo Jeong, Rachata Ausavarungnirun 等USENIX ATC 2022 · 被引用 28 次
- CHOPIN: Scalable Graphics Rendering in Multi-GPU Systems via Parallel Image CompositionXiaowei Ren, Mieszko LisHPCA 2021 · 被引用 14 次
