NearFetch: Saving Inter-Module Bandwidth in Many-Chip-Module GPUs
Xia Zhao, Guangda Zhang, Lu Wang, Shiqing Zhang, Huadong Dai
Abstract
As Graphics Processing Units (GPUs) face increasing computing demands that surpass single-module capabilities due to transistor scaling and lithography constraints, the necessity for expanding the module count within GPUs grows. This escalation faces a significant challenge: the total inter-module bandwidth in many-chip-module GPUs is limited by manufacturing constraints in organic substrates or silicon interposers. Unlike Central Processing Units (CPUs), which are latency-sensitive, GPUs leverage their high thread-level parallelism to effectively hide memory access latency through simultaneous multithreading. This attribute makes GPUs inherently sensitive to bandwidth constraints, making the efficient exploitation of available inter-module bandwidth important. In this paper, we identify that fetching data from faraway memory in many-chip-module GPUs can easily cause bandwidth contention which degrades the real achieved data bandwidth per GPU module compared to fetching data from nearby memory. To further analyze this problem, we introduce the Inter-Module Bandwidth per Access (IBPA) metric for quantifying bandwidth usage and finding the network hop count directly impacts the IBPA and network contention. Next, we propose NearFetch, a routing-based solution to reduce IBPA. NearFetch works due to the fact that GPU modules along the routing path are typically much closer to the source GPU module while these GPU modules can supply of data for high-sharing applications. NearFetch consists of two primary components: a data forwarding scheme, enabling data forwarding when the data resides in a remote GPU module, and a topology-aware Miss Status Handling Register (MSHR) coalescing scheme, responsible for recording the memory address information for future use in case of a data miss. By leveraging the data locality among various GPU modules, NearFetch substantially minimizes inter-module bandwidth usage, eliminating the need to fetch data from distant memory partitions. Our evaluation of NearFetch within the context of many-chip-module GPUs, across applications exhibiting diverse degrees of data locality, reveals that it reduces IBPA by and enhances performance by an average of (with up to improvement) for high-sharing workloads.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 6bc28f64-3d93-4108-b3c1-7bee8ec4def4Related papers
- Snake: A Variable-length Chain-based Prefetching for GPUsSaba Mostofi, Hajar Falahati, Negin Mahani, Pejman Lotfi-Kamran et al.MICRO 2023 · 11 citations
- The Sparsity-Aware LazyGPU ArchitectureChangxi Liu, Miao Yu, Yifan Sun, Trevor E. CarlsonISCA 2025 · 1 citation
- HDPAT: Hierarchical Distributed Page Address Translation for Wafer-Scale GPUsDaoxuan Xu, Ying Li, Yuwei Sun, Jie Ren et al.HPCA 2026 · 1 citation
- Griffin: Hardware-Software Support for Efficient Page Migration in Multi-GPU SystemsTrinayan Baruah, Yifan Sun, Ali Tolga Dinçer, Saiful A. Mojumder et al.HPCA 2020 · 50 citations
- NetCrafter: Tailoring Network Traffic for Non-Uniform Bandwidth Multi-GPU SystemsAmel Fatima, Yang Yang, Yifan Sun, Rachata Ausavarungnirun et al.ISCA 2025 · 2 citations
