Principal Kernel Analysis: A Tractable Methodology to Simulate Scaled GPU Workloads
Cesar Avalos Baddouh, Mahmoud Khairy, Roland N. Green, Mathias Payer, Timothy G. Rogers
Abstract
Simulating all threads in a scaled GPU workload results in prohibitive simulation cost. Cycle-level simulation is orders of magnitude slower than native silicon, the only solution is to reduce the amount of work simulated while accurately representing the program.
Existing solutions to simulate GPU programs either scale the input size, simulate the first several billion instructions, or simulate a portion of both the GPU and the workload. These solutions lack validation against scaled systems, produce unrealistic contention conditions and frequently miss critical code sections. Existing CPU sampling mechanisms, like SimPoint, reduce per-thread workload, and are ill-suited to GPU programs where reducing the number of threads is critical. Sampling solutions on GPUs space lack silicon validation, require per-workload parameter tuning, and do not scale.
A tractable solution, validated on contemporary scaled workloads, is needed to provide credible simulation results. By studying scaled workloads with centuries-long simulation times, we uncover practical and algorithmic limitations of existing solutions and propose Principal Kernel Analysis: a hierarchical program sampling methodology that concisely represents GPU programs by selecting representative kernel portions using a scalable profiling methodology, tractable clustering algorithm and detection of intra-kernel IPC stability. We validate Principal Kernel Analysis across 147 workloads and three GPU generations using the Accel-Sim simulator, demonstrating a better performance/error tradeoff than prior work and that century-long MLPerf simulations are reduced to hours with an average cycle error of 27% versus silicon.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 69ee58ca-99c5-4ff9-96bd-ccb849948c00Cited by top-tier papers8
- Forecasting GPU Performance for Deep Learning Training and InferenceSeonho Lee, Amar Phanishayee, Divya MahajanASPLOS 2025 · 31 citations
- GPU Scale-Model SimulationHossein SeyyedAghaei, Mahmood Naderan-Tahan, Lieven EeckhoutHPCA 2024 · 13 citations
- Treelet Prefetching For Ray TracingYuan-Hsi Chou, Tyler Nowicki, Tor M. AamodtMICRO 2023 · 12 citations
- Photon: A Fine-grained Sampled Simulation Methodology for GPU WorkloadsChangxi Liu, Yifan Sun, Trevor E. CarlsonMICRO 2023 · 10 citations
- Swift and Trustworthy Large-Scale GPU Simulation with Fine-Grained Error Modeling and Hierarchical ClusteringEuijun Chung, Seonjin Na, Sung Ha Kang, Hyesoon KimMICRO 2025 · 5 citations
Builds on2
Related papers
- Scalable Deep Learning-Based Microarchitecture Simulation on GPUsSantosh Pandey, Lingda Li, Thomas Flynn, Adolfy Hoisie et al.SC 2022 · 7 citations
- GEM: GPU-Accelerated Emulator-Inspired RTL SimulationZizheng Guo, Yanqing Zhang, Runsheng Wang, Yibo Lin et al.DAC 2025 · 3 citations
- LoopPoint: Checkpoint-driven Sampled Simulation for Multi-threaded ApplicationsAlen Sabu, Harish Patil, Wim Heirman, Trevor E. CarlsonHPCA 2022 · 21 citations
- GCStack+GCScaler: Fast and Accurate GPU Performance Analyses Using Fine-Grained Stall Cycle Accounting and Interval AnalysisHanna Cha, Sungchul Lee, Jounghoo Lee, Yeonan Ha et al.ISCA 2025 · 1 citation
- ThreadFuser: A SIMT Analysis Framework for MIMD ProgramsAhmad Alawneh, Ni Kang, Mahmoud Khairy, Timothy G. RogersMICRO 2024
