Lune

SC2022顶会

Climbing the Summit and Pushing the Frontier of Mixed Precision Benchmarks at Extreme Scale

Hao Lu, Michael A. Matheson, Vladyslav Oles, J. Austin Ellis, Wayne Joubert, Feiyi Wang

2022年份
8被引次数

摘要

accuracy. The convergence of HPC and ML offers the opportunity to develop new techniques that exploit mixed precision capabilities to enable new science. The High Performance LINPACK (HPL) benchmark [2] has long played a key role in tracking the performance of the world's top supercomputers. The double precision math reflects the precision used in typical HPC applications. Most of the ML models seen in centers train using mixed precision operations, which are fundamentally different from those in HPL.

The High-Performance LINPACK benchmark for Accelerator Inspection (HPL-AI) [3] was created to provide a benchmark that measures the performance potential of mixed precision (FP16/FP32) on newly deployed supercomputers. Together these benchmarks are important in the acquisition of large leadership class supercomputers given the cost and resources necessary to design, deploy, and maintain these systems. The combined benchmarks yield valuable insight into the double precision and mixed precision performance possible on systems covering a wide range of use cases.

Only a handful of systems in the world are capable of exascale performance, and just a limited number of applications are able to sustain it. The present work helps pave the way for more into the future. The main contributions of this paper are:

• We summarize the efforts and accomplishments of developing a leadership class cross-platform implementation of the HPL-AI benchmark for two of the world's fastest supercomputers at the OLCF, both the NVIDIA-based Summit and the newly launched AMD-based exascale system, Frontier. This is the first known code that runs on different GPU enabled systems and delivers exascale performance on both. We report sustained performance of 1.411 EFLOPS on Summit and 2.387 EFLOPS on approximately 40% of Frontier. Our Summit result achieved 9.5 times the performance of HPL demonstrating the value of mixed precision.

• We propose a performance model based on measured floating point operation (flop) rate for key compute ker-Abstract-The rise of machine learning (ML) applications and their use of mixed precision to perform interesting science are driving forces behind AI for science on HPC. The convergence of ML and HPC with mixed precision offers the possibility of transformational changes in computational science. The HPL-AI benchmark is designed to measure the performance of mixed precision arithmetic as opposed to the HPL benchmark which measures double precision performance. Pushing the limits of systems at extreme scale is nontrivial -little public literature explores optimization of mixed precision computations at this scale. In this work, we demonstrate how to scale up the HPL-AI benchmark on the pre-exascale Summit and exascale Frontier systems at the Oak Ridge Leadership Computing Facility (OLCF) with a cross-platform design. We present the implementation, performance results, and a guideline of optimization strategies employed for delivering portable performance on both AMD and NVIDIA GPUs at extreme scale.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper1

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖