Dissecting and Modeling the Architecture of Modern GPU Cores
Rodrigo Huerta, Mojtaba Abaie Shoushtary, José-Lorenzo Cruz, Antonio González
摘要
GPUs are the most popular platform for accelerating HPC workloads, such as artificial intelligence and science simulations. However, most microarchitectural research in academia relies on GPU core pipeline designs based on architectures that are more than 15 years old.
This paper reverse engineers modern NVIDIA GPU cores, unveiling many key aspects of its design and explaining how GPUs leverage hardware-compiler techniques where the compiler guides hardware during execution. In particular, it reveals how the issue logic works including the policy of the issue scheduler, the structure of the register file and its associated cache, and multiple features of the memory pipeline. Moreover, it analyses how a simple instruction prefetcher based on a stream buffer fits well with modern NVIDIA GPUs and is likely to be used. Furthermore, we investigate the impact of the register file cache and the number of register file read ports on both simulation accuracy and performance.
By modeling all these new discovered microarchitectural details, we achieve 18.24% lower mean absolute percentage error (MAPE) in execution cycles than previous state-of-the-art simulators, resulting in an average of 13.98% MAPE with respect to real hardware (NVIDIA RTX A6000). Also, we demonstrate that this new model stands for other NVIDIA architectures, such as Turing.
Finally, we show that the software-based dependence management mechanism included in modern NVIDIA GPUs outperforms a hardware mechanism based on scoreboards in terms of performance and area.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Swift and Trustworthy Large-Scale GPU Simulation with Fine-Grained Error Modeling and Hierarchical ClusteringEuijun Chung, Seonjin Na, Sung Ha Kang, Hyesoon KimMICRO 2025 · 被引用 5 次
- Fuzzing Open-Source GPU Hardware with SIMT Program GenerationZibo Gao, Jie Wang, Qihang Zhou, Lixiao Shan 等USENIX Security 2026
- BCCE: Block-Centric GPU Co-Design for Real-Time Range-Top-K Query at ScaleChengying Huan, Ziheng Meng, Zhengyi Yang, Yongchao Liu 等HPDC 2026
它引用的顶会 Paper8
- Accel-Sim: An Extensible Simulation Framework for Validated GPU ModelingMahmoud Khairy, Zhesheng Shen, Tor M. Aamodt, Timothy G. RogersISCA 2020 · 被引用 366 次
- Network-on-Chip Microarchitecture-based Covert Channel in GPUsJaeguk Ahn, Jiho Kim, Hans Kasan, Zhixian Jin 等MICRO 2021 · 被引用 30 次
- BOW: Breathing Operand Windows to Exploit Bypassing in GPUsHodjat Asghari Esfeden, AmirAli Abdolrashidi, Shafiur Rahman, Daniel Wong 等MICRO 2020 · 被引用 19 次
- Mitigating GPU Core Partitioning Performance EffectsAaron Barnes, Fangjia Shen, Timothy G. RogersHPCA 2023 · 被引用 18 次
- TunneLs for Bootlegging: Fully Reverse-Engineering GPU TLBs for Challenging Isolation Guarantees of NVIDIA MIGZhenkai Zhang, Tyler N. Allen, Fan Yao, Xing Gao 等CCS 2023 · 被引用 18 次
相关 Paper
- GCoM: a detailed GPU core model for accurate analytical modeling of modern GPUsJounghoo Lee, Yeonan Ha, Suhyun Lee, Jinyoung Woo 等ISCA 2022 · 被引用 25 次
- WASP: Exploiting GPU Pipeline Parallelism with Hardware-Accelerated Automatic Warp SpecializationNeal Clayton Crago, Sana Damani, Karthikeyan Sankaralingam, Stephen W. KecklerHPCA 2024 · 被引用 13 次
- Demystifying and Exploiting ASLR on NVIDIA GPUsRuofan Zhu, Ganhao Chen, Wenbo Shen, Lyuye Zhang 等S&P 2026 · 被引用 2 次
- Scalable Deep Learning-Based Microarchitecture Simulation on GPUsSantosh Pandey, Lingda Li, Thomas Flynn, Adolfy Hoisie 等SC 2022 · 被引用 7 次
- GCStack+GCScaler: Fast and Accurate GPU Performance Analyses Using Fine-Grained Stall Cycle Accounting and Interval AnalysisHanna Cha, Sungchul Lee, Jounghoo Lee, Yeonan Ha 等ISCA 2025 · 被引用 1 次
