Principal Kernel Analysis: A Tractable Methodology to Simulate Scaled GPU Workloads
Cesar Avalos Baddouh, Mahmoud Khairy, Roland N. Green, Mathias Payer, Timothy G. Rogers
摘要
Simulating all threads in a scaled GPU workload results in prohibitive simulation cost. Cycle-level simulation is orders of magnitude slower than native silicon, the only solution is to reduce the amount of work simulated while accurately representing the program.
Existing solutions to simulate GPU programs either scale the input size, simulate the first several billion instructions, or simulate a portion of both the GPU and the workload. These solutions lack validation against scaled systems, produce unrealistic contention conditions and frequently miss critical code sections. Existing CPU sampling mechanisms, like SimPoint, reduce per-thread workload, and are ill-suited to GPU programs where reducing the number of threads is critical. Sampling solutions on GPUs space lack silicon validation, require per-workload parameter tuning, and do not scale.
A tractable solution, validated on contemporary scaled workloads, is needed to provide credible simulation results. By studying scaled workloads with centuries-long simulation times, we uncover practical and algorithmic limitations of existing solutions and propose Principal Kernel Analysis: a hierarchical program sampling methodology that concisely represents GPU programs by selecting representative kernel portions using a scalable profiling methodology, tractable clustering algorithm and detection of intra-kernel IPC stability. We validate Principal Kernel Analysis across 147 workloads and three GPU generations using the Accel-Sim simulator, demonstrating a better performance/error tradeoff than prior work and that century-long MLPerf simulations are reduced to hours with an average cycle error of 27% versus silicon.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Forecasting GPU Performance for Deep Learning Training and InferenceSeonho Lee, Amar Phanishayee, Divya MahajanASPLOS 2025 · 被引用 31 次
- GPU Scale-Model SimulationHossein SeyyedAghaei, Mahmood Naderan-Tahan, Lieven EeckhoutHPCA 2024 · 被引用 13 次
- Treelet Prefetching For Ray TracingYuan-Hsi Chou, Tyler Nowicki, Tor M. AamodtMICRO 2023 · 被引用 12 次
- Photon: A Fine-grained Sampled Simulation Methodology for GPU WorkloadsChangxi Liu, Yifan Sun, Trevor E. CarlsonMICRO 2023 · 被引用 10 次
- Swift and Trustworthy Large-Scale GPU Simulation with Fine-Grained Error Modeling and Hierarchical ClusteringEuijun Chung, Seonjin Na, Sung Ha Kang, Hyesoon KimMICRO 2025 · 被引用 5 次
它引用的顶会 Paper2
相关 Paper
- Scalable Deep Learning-Based Microarchitecture Simulation on GPUsSantosh Pandey, Lingda Li, Thomas Flynn, Adolfy Hoisie 等SC 2022 · 被引用 7 次
- GEM: GPU-Accelerated Emulator-Inspired RTL SimulationZizheng Guo, Yanqing Zhang, Runsheng Wang, Yibo Lin 等DAC 2025 · 被引用 3 次
- LoopPoint: Checkpoint-driven Sampled Simulation for Multi-threaded ApplicationsAlen Sabu, Harish Patil, Wim Heirman, Trevor E. CarlsonHPCA 2022 · 被引用 21 次
- GCStack+GCScaler: Fast and Accurate GPU Performance Analyses Using Fine-Grained Stall Cycle Accounting and Interval AnalysisHanna Cha, Sungchul Lee, Jounghoo Lee, Yeonan Ha 等ISCA 2025 · 被引用 1 次
- ThreadFuser: A SIMT Analysis Framework for MIMD ProgramsAhmad Alawneh, Ni Kang, Mahmoud Khairy, Timothy G. RogersMICRO 2024
