Swift and Trustworthy Large-Scale GPU Simulation with Fine-Grained Error Modeling and Hierarchical Clustering
Euijun Chung, Seonjin Na, Sung Ha Kang, Hyesoon Kim
Abstract
Kernel-level sampling is an effective technique for running largescale GPU workloads on cycle-level simulators by selecting a representative subset of kernels, thereby significantly reducing simulation complexity and runtime. However, in large-scale GPU workloads, kernels often exhibit heterogeneous runtime behaviors where some identical kernels show fluctuating performance, while others display multiple performance saturation points. We observe that the kernel execution time distribution is a powerful signature for addressing this complexity. By carefully analyzing execution time distributions, we show that heterogeneous kernels can be effectively classified and sampled, significantly reducing errors in sampled simulations.
This paper proposes STEM+ROOT, a fine-grained kernel-level sampling methodology that enables trustworthy sampled simulation by achieving minimal sampling error. STEM leverages the distribution of kernel execution times as a signature and applies statistical techniques to determine optimal sample sizes with tight error bounds. ROOT is a novel hierarchical clustering framework built on top of STEM that ensures the sampled kernels faithfully represent the entire workload in terms of execution time and a wide range of microarchitectural metrics. STEM achieves high scalability for large-scale GPU workloads by significantly reducing offline profiling overhead for collecting kernel execution times. When evaluated on the latest GPU benchmark suite, our proposed methodology reduces sampling error by 27.6-81.9× and achieves 7-600× faster kernel profiling than existing approaches while achieving comparable performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2a7f0639-a069-45c0-9caa-10f779034534Builds on9
- Accel-Sim: An Extensible Simulation Framework for Validated GPU ModelingMahmoud Khairy, Zhesheng Shen, Tor M. Aamodt, Timothy G. RogersISCA 2020 · 366 citations
- Principal Kernel Analysis: A Tractable Methodology to Simulate Scaled GPU WorkloadsCesar Avalos Baddouh, Mahmoud Khairy, Roland N. Green, Mathias Payer et al.MICRO 2021 · 26 citations
- LoopPoint: Checkpoint-driven Sampled Simulation for Multi-threaded ApplicationsAlen Sabu, Harish Patil, Wim Heirman, Trevor E. CarlsonHPCA 2022 · 21 citations
- GPU Scale-Model SimulationHossein SeyyedAghaei, Mahmood Naderan-Tahan, Lieven EeckhoutHPCA 2024 · 13 citations
- Photon: A Fine-grained Sampled Simulation Methodology for GPU WorkloadsChangxi Liu, Yifan Sun, Trevor E. CarlsonMICRO 2023 · 10 citations
Related papers
- HyFiSS: A Hybrid Fidelity Stall-Aware Simulator for GPGPUsJianchao Yang, Mei Wen, Dong Chen, Zhaoyun Chen et al.MICRO 2024 · 8 citations
- GVARP: Detecting Performance Variance on Large-Scale Heterogeneous SystemsXin You, Zhibo Xuan, Hailong Yang, Zhongzhi Luan et al.SC 2024 · 8 citations
- HeteroSim: Towards High-Fidelity Heterogeneous LLM Training Simulation on GPUsXiaofei Yue, Fangming Zhao, Fulun Ye, Jiongchi Yu et al.WWW 2026
- Exploration of LLM Workload Reliability Based on di/dt Effects and Voltage DroopsZhixing Jiang, Justin Garrigus, Allison Seigler, Ethan Syed et al.HPCA 2026
- Nugget: Portable Program SnippetsZhantong Qiu, Mahyar Samani, Jason Lowe-PowerHPCA 2026
