Uncovering Real GPU NoC Characteristics: Implications on Interconnect Architecture
Zhixian Jin, Christopher Rocca, Jiho Kim, Hans Kasan, Minsoo Rhu, Ali Bakhoda, Tor M. Aamodt, John Kim
Abstract
A critical component of high-throughput processors such as GPUs is the network-on-chip (NoC) that interconnects the large number of cores and the memory partitions together. In this work, we provide a detailed analysis, in terms of latency and bandwidth, of real GPU NoC across several generations of modern NVIDIA GPUs. Our analysis identifies how non-uniform latency exists between the cores and the memory partitions based on their physical location in the GPU. The non-uniformity can result in up to approximately 70 % difference in on-chip latency. In comparison, the bandwidth provided from the cores to the memory partitions is approximately uniform. However, recent GPUs that consist of multiple GPU “partitions” present different on-chip latency and bandwidth characteristics when communicating between the partitions. Based on our analysis of real GPU interconnect, we discuss potential implications including its impact on timing used in side-channel attacks as well as NoC microarchitectures. We show how the non-uniform latency can be exploited in a timing side-channel attack within a GPU as the core location impacts performance (or timing). In addition, proper understanding (and proper assumptions) of GPU NoC is critical to ensure a network that does not bottleneck the overall system performance.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get cba2e369-a7ec-4653-af0e-397db72e7603Cited by top-tier papers6
- ClusterFusion: Expanding Operator Fusion Scope for LLM Inference via Cluster-Level Collective PrimitiveXinhao Luo, Zihan Liu, Yangjie Zhou, Shihan Fang et al.NeurIPS 2025 · 9 citations
- Dissecting and Modeling the Architecture of Modern GPU CoresRodrigo Huerta, Mojtaba Abaie Shoushtary, José-Lorenzo Cruz, Antonio GonzálezMICRO 2025 · 8 citations
- Characterizing the Efficiency of Distributed Training: A Power, Performance, and Thermal PerspectiveSeokjin Go, Joongun Park, Spandan More, Hanjiang Wu et al.MICRO 2025 · 7 citations
- Exploiting TLBs in Virtualized GPUs for Cross-VM Side-Channel AttacksHongyue Jin, Yanan Guo, Zhenkai ZhangNDSS 2026
- FlashFuser: Expanding the Scale of Kernel Fusion for Compute-Intensive Operators via Inter-Core ConnectionZiyu Huang, Yangjie Zhou, Zihan Liu, Xinhao Luo et al.HPCA 2026
Related papers
- Network-on-Chip Microarchitecture-based Covert Channel in GPUsJaeguk Ahn, Jiho Kim, Hans Kasan, Zhixian Jin et al.MICRO 2021 · 30 citations
- Ghost Arbitration: Mitigating Interconnect Side-Channel Timing Attacks in GPUZhixian Jin, Jaeguk Ahn, Jiho Kim, Hans Kasan et al.MICRO 2024 · 4 citations
- Veiled Pathways: Investigating Covert and Side Channels Within GPU UncoreYuanqing Miao, Yingtian Zhang, Dinghao Wu, Danfeng Zhang et al.MICRO 2024 · 8 citations
- Trident: A Hybrid Correlation-Collision GPU Cache Timing Attack for AES Key RecoveryJaeguk Ahn, Cheolgyu Jin, Jiho Kim, Minsoo Rhu et al.HPCA 2021 · 15 citations
- TIMESLICE-SANDWICH: A GPU Side-Channel Attack Exploiting Time-Sliced SchedulingHodong Kim, Gyeongsup Lim, Seunghee Shin, Youngjoo Shin et al.USENIX Security 2026
