SC2020Top-tier venue
GVProf: a value profiler for GPU-based clusters
Keren Zhou, Yueming Hao, John M. Mellor-Crummey, Xiaozhu Meng, Xu Liu
Abstract
GPGPUs are widely used in high-performance computing systems to accelerate scientific and machine learning workloads. Developing efficient GPU kernels is critically important to obtain “bare-metal” performance on GPU-based clusters. In this paper, we describe the design and implementation of GVPROF, the first value profiler that pinpoints value-related inefficiencies in applications running on NVIDIA GPU-based clusters. The novelty of GVPROF resides in its ability to detect temporal and spatial value redundancies, which provides useful information to guide code optimization. GVPROF can monitor production multi-node multi-GPU executions in clusters. Our experiments with well-known GPU benchmarks and HPC applications show that GVPROF incurs acceptable overhead and scales to large executions. Using GVPROF, we optimized several HPC and machine learning workloads on one NVIDIA V100 GPU. In one case study of LAMMPS, optimizations based on information from GVProf led to whole-program speedups ranging from 1.37x on a single GPU to 1.08x on 64 GPUs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 30991f26-96bf-4750-ad8c-e4f3cadbd831Cited by top-tier papers3
- GREYHOUND: Hunting Fail-Slows in Hybrid-Parallel Training at ScaleTianyuan Wu, Wei Wang, Yinghao Yu, Siran Yang et al.USENIX ATC 2025 · 19 citations
- TrioSim: A Lightweight Simulator for Large-Scale DNN Workloads on Multi-GPU SystemsYing Li, Yuhui Bao, Gongyu Wang, Xinxin Mei et al.ISCA 2025 · 2 citations
- KPerfIR: Towards a Open and Compiler-centric Ecosystem for GPU Kernel Performance Tooling on Modern AI WorkloadsYue Guan, Yuanwei Fang, Keren Zhou, Corbin Robeck et al.OSDI 2025
Builds on1
Related papers
- ValueExpert: exploring value patterns in GPU-accelerated applicationsKeren Zhou, Yueming Hao, John M. Mellor-Crummey, Xiaozhu Meng et al.ASPLOS 2022 · 19 citations
- DrGPUM: Guiding Memory Optimization for GPU-Accelerated ApplicationsMao Lin, Keren Zhou, Pengfei SuASPLOS 2023 · 13 citations
- RedSan: A Redundant Memory Instruction Sanitizer for GPU ProgramsYanbo Zhao, Yueming Hao, Zecheng Li, Shuyin Jiao et al.SC 2025 · 1 citation
- Improving GPU Energy Efficiency through an Application-transparent Frequency Scaling Policy with Performance AssuranceYijia Zhang, Qiang Wang, Zhe Lin, Pengxiang Xu et al.EuroSys 2024 · 18 citations
- PAL: A Variability-Aware Policy for Scheduling ML Workloads in GPU ClustersRutwik Jain, Brandon Tran, Keting Chen, Matthew D. Sinclair et al.SC 2024 · 10 citations
