Uncovering Real GPU NoC Characteristics: Implications on Interconnect Architecture
Zhixian Jin, Christopher Rocca, Jiho Kim, Hans Kasan, Minsoo Rhu, Ali Bakhoda, Tor M. Aamodt, John Kim
摘要
A critical component of high-throughput processors such as GPUs is the network-on-chip (NoC) that interconnects the large number of cores and the memory partitions together. In this work, we provide a detailed analysis, in terms of latency and bandwidth, of real GPU NoC across several generations of modern NVIDIA GPUs. Our analysis identifies how non-uniform latency exists between the cores and the memory partitions based on their physical location in the GPU. The non-uniformity can result in up to approximately 70 % difference in on-chip latency. In comparison, the bandwidth provided from the cores to the memory partitions is approximately uniform. However, recent GPUs that consist of multiple GPU “partitions” present different on-chip latency and bandwidth characteristics when communicating between the partitions. Based on our analysis of real GPU interconnect, we discuss potential implications including its impact on timing used in side-channel attacks as well as NoC microarchitectures. We show how the non-uniform latency can be exploited in a timing side-channel attack within a GPU as the core location impacts performance (or timing). In addition, proper understanding (and proper assumptions) of GPU NoC is critical to ensure a network that does not bottleneck the overall system performance.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper6
- ClusterFusion: Expanding Operator Fusion Scope for LLM Inference via Cluster-Level Collective PrimitiveXinhao Luo, Zihan Liu, Yangjie Zhou, Shihan Fang 等NeurIPS 2025 · 被引用 9 次
- Dissecting and Modeling the Architecture of Modern GPU CoresRodrigo Huerta, Mojtaba Abaie Shoushtary, José-Lorenzo Cruz, Antonio GonzálezMICRO 2025 · 被引用 8 次
- Characterizing the Efficiency of Distributed Training: A Power, Performance, and Thermal PerspectiveSeokjin Go, Joongun Park, Spandan More, Hanjiang Wu 等MICRO 2025 · 被引用 7 次
- Exploiting TLBs in Virtualized GPUs for Cross-VM Side-Channel AttacksHongyue Jin, Yanan Guo, Zhenkai ZhangNDSS 2026
- FlashFuser: Expanding the Scale of Kernel Fusion for Compute-Intensive Operators via Inter-Core ConnectionZiyu Huang, Yangjie Zhou, Zihan Liu, Xinhao Luo 等HPCA 2026
相关 Paper
- Network-on-Chip Microarchitecture-based Covert Channel in GPUsJaeguk Ahn, Jiho Kim, Hans Kasan, Zhixian Jin 等MICRO 2021 · 被引用 30 次
- Ghost Arbitration: Mitigating Interconnect Side-Channel Timing Attacks in GPUZhixian Jin, Jaeguk Ahn, Jiho Kim, Hans Kasan 等MICRO 2024 · 被引用 4 次
- Veiled Pathways: Investigating Covert and Side Channels Within GPU UncoreYuanqing Miao, Yingtian Zhang, Dinghao Wu, Danfeng Zhang 等MICRO 2024 · 被引用 8 次
- Trident: A Hybrid Correlation-Collision GPU Cache Timing Attack for AES Key RecoveryJaeguk Ahn, Cheolgyu Jin, Jiho Kim, Minsoo Rhu 等HPCA 2021 · 被引用 15 次
- TIMESLICE-SANDWICH: A GPU Side-Channel Attack Exploiting Time-Sliced SchedulingHodong Kim, Gyeongsup Lim, Seunghee Shin, Youngjoo Shin 等USENIX Security 2026
