GVProf: a value profiler for GPU-based clusters
Keren Zhou, Yueming Hao, John M. Mellor-Crummey, Xiaozhu Meng, Xu Liu
摘要
GPGPUs are widely used in high-performance computing systems to accelerate scientific and machine learning workloads. Developing efficient GPU kernels is critically important to obtain “bare-metal” performance on GPU-based clusters. In this paper, we describe the design and implementation of GVPROF, the first value profiler that pinpoints value-related inefficiencies in applications running on NVIDIA GPU-based clusters. The novelty of GVPROF resides in its ability to detect temporal and spatial value redundancies, which provides useful information to guide code optimization. GVPROF can monitor production multi-node multi-GPU executions in clusters. Our experiments with well-known GPU benchmarks and HPC applications show that GVPROF incurs acceptable overhead and scales to large executions. Using GVPROF, we optimized several HPC and machine learning workloads on one NVIDIA V100 GPU. In one case study of LAMMPS, optimizations based on information from GVProf led to whole-program speedups ranging from 1.37x on a single GPU to 1.08x on 64 GPUs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- GREYHOUND: Hunting Fail-Slows in Hybrid-Parallel Training at ScaleTianyuan Wu, Wei Wang, Yinghao Yu, Siran Yang 等USENIX ATC 2025 · 被引用 19 次
- TrioSim: A Lightweight Simulator for Large-Scale DNN Workloads on Multi-GPU SystemsYing Li, Yuhui Bao, Gongyu Wang, Xinxin Mei 等ISCA 2025 · 被引用 2 次
- KPerfIR: Towards a Open and Compiler-centric Ecosystem for GPU Kernel Performance Tooling on Modern AI WorkloadsYue Guan, Yuanwei Fang, Keren Zhou, Corbin Robeck 等OSDI 2025
它引用的顶会 Paper1
相关 Paper
- ValueExpert: exploring value patterns in GPU-accelerated applicationsKeren Zhou, Yueming Hao, John M. Mellor-Crummey, Xiaozhu Meng 等ASPLOS 2022 · 被引用 19 次
- DrGPUM: Guiding Memory Optimization for GPU-Accelerated ApplicationsMao Lin, Keren Zhou, Pengfei SuASPLOS 2023 · 被引用 13 次
- RedSan: A Redundant Memory Instruction Sanitizer for GPU ProgramsYanbo Zhao, Yueming Hao, Zecheng Li, Shuyin Jiao 等SC 2025 · 被引用 1 次
- Improving GPU Energy Efficiency through an Application-transparent Frequency Scaling Policy with Performance AssuranceYijia Zhang, Qiang Wang, Zhe Lin, Pengxiang Xu 等EuroSys 2024 · 被引用 18 次
- PAL: A Variability-Aware Policy for Scheduling ML Workloads in GPU ClustersRutwik Jain, Brandon Tran, Keting Chen, Matthew D. Sinclair 等SC 2024 · 被引用 10 次
