Swift and Trustworthy Large-Scale GPU Simulation with Fine-Grained Error Modeling and Hierarchical Clustering
Euijun Chung, Seonjin Na, Sung Ha Kang, Hyesoon Kim
摘要
Kernel-level sampling is an effective technique for running largescale GPU workloads on cycle-level simulators by selecting a representative subset of kernels, thereby significantly reducing simulation complexity and runtime. However, in large-scale GPU workloads, kernels often exhibit heterogeneous runtime behaviors where some identical kernels show fluctuating performance, while others display multiple performance saturation points. We observe that the kernel execution time distribution is a powerful signature for addressing this complexity. By carefully analyzing execution time distributions, we show that heterogeneous kernels can be effectively classified and sampled, significantly reducing errors in sampled simulations.
This paper proposes STEM+ROOT, a fine-grained kernel-level sampling methodology that enables trustworthy sampled simulation by achieving minimal sampling error. STEM leverages the distribution of kernel execution times as a signature and applies statistical techniques to determine optimal sample sizes with tight error bounds. ROOT is a novel hierarchical clustering framework built on top of STEM that ensures the sampled kernels faithfully represent the entire workload in terms of execution time and a wide range of microarchitectural metrics. STEM achieves high scalability for large-scale GPU workloads by significantly reducing offline profiling overhead for collecting kernel execution times. When evaluated on the latest GPU benchmark suite, our proposed methodology reduces sampling error by 27.6-81.9× and achieves 7-600× faster kernel profiling than existing approaches while achieving comparable performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- Accel-Sim: An Extensible Simulation Framework for Validated GPU ModelingMahmoud Khairy, Zhesheng Shen, Tor M. Aamodt, Timothy G. RogersISCA 2020 · 被引用 366 次
- Principal Kernel Analysis: A Tractable Methodology to Simulate Scaled GPU WorkloadsCesar Avalos Baddouh, Mahmoud Khairy, Roland N. Green, Mathias Payer 等MICRO 2021 · 被引用 26 次
- LoopPoint: Checkpoint-driven Sampled Simulation for Multi-threaded ApplicationsAlen Sabu, Harish Patil, Wim Heirman, Trevor E. CarlsonHPCA 2022 · 被引用 21 次
- GPU Scale-Model SimulationHossein SeyyedAghaei, Mahmood Naderan-Tahan, Lieven EeckhoutHPCA 2024 · 被引用 13 次
- Photon: A Fine-grained Sampled Simulation Methodology for GPU WorkloadsChangxi Liu, Yifan Sun, Trevor E. CarlsonMICRO 2023 · 被引用 10 次
相关 Paper
- HyFiSS: A Hybrid Fidelity Stall-Aware Simulator for GPGPUsJianchao Yang, Mei Wen, Dong Chen, Zhaoyun Chen 等MICRO 2024 · 被引用 8 次
- GVARP: Detecting Performance Variance on Large-Scale Heterogeneous SystemsXin You, Zhibo Xuan, Hailong Yang, Zhongzhi Luan 等SC 2024 · 被引用 8 次
- HeteroSim: Towards High-Fidelity Heterogeneous LLM Training Simulation on GPUsXiaofei Yue, Fangming Zhao, Fulun Ye, Jiongchi Yu 等WWW 2026
- Exploration of LLM Workload Reliability Based on di/dt Effects and Voltage DroopsZhixing Jiang, Justin Garrigus, Allison Seigler, Ethan Syed 等HPCA 2026
- Nugget: Portable Program SnippetsZhantong Qiu, Mahyar Samani, Jason Lowe-PowerHPCA 2026
