Lune

SC2022Top-tier venue

Climbing the Summit and Pushing the Frontier of Mixed Precision Benchmarks at Extreme Scale

Hao Lu, Michael A. Matheson, Vladyslav Oles, J. Austin Ellis, Wayne Joubert, Feiyi Wang

2022Year
8Citations

Abstract

accuracy. The convergence of HPC and ML offers the opportunity to develop new techniques that exploit mixed precision capabilities to enable new science. The High Performance LINPACK (HPL) benchmark [2] has long played a key role in tracking the performance of the world's top supercomputers. The double precision math reflects the precision used in typical HPC applications. Most of the ML models seen in centers train using mixed precision operations, which are fundamentally different from those in HPL.

The High-Performance LINPACK benchmark for Accelerator Inspection (HPL-AI) [3] was created to provide a benchmark that measures the performance potential of mixed precision (FP16/FP32) on newly deployed supercomputers. Together these benchmarks are important in the acquisition of large leadership class supercomputers given the cost and resources necessary to design, deploy, and maintain these systems. The combined benchmarks yield valuable insight into the double precision and mixed precision performance possible on systems covering a wide range of use cases.

Only a handful of systems in the world are capable of exascale performance, and just a limited number of applications are able to sustain it. The present work helps pave the way for more into the future. The main contributions of this paper are:

• We summarize the efforts and accomplishments of developing a leadership class cross-platform implementation of the HPL-AI benchmark for two of the world's fastest supercomputers at the OLCF, both the NVIDIA-based Summit and the newly launched AMD-based exascale system, Frontier. This is the first known code that runs on different GPU enabled systems and delivers exascale performance on both. We report sustained performance of 1.411 EFLOPS on Summit and 2.387 EFLOPS on approximately 40% of Frontier. Our Summit result achieved 9.5 times the performance of HPL demonstrating the value of mixed precision.

• We propose a performance model based on measured floating point operation (flop) rate for key compute ker-Abstract-The rise of machine learning (ML) applications and their use of mixed precision to perform interesting science are driving forces behind AI for science on HPC. The convergence of ML and HPC with mixed precision offers the possibility of transformational changes in computational science. The HPL-AI benchmark is designed to measure the performance of mixed precision arithmetic as opposed to the HPL benchmark which measures double precision performance. Pushing the limits of systems at extreme scale is nontrivial -little public literature explores optimization of mixed precision computations at this scale. In this work, we demonstrate how to scale up the HPL-AI benchmark on the pre-exascale Summit and exascale Frontier systems at the Oak Ridge Leadership Computing Facility (OLCF) with a cross-platform design. We present the implementation, performance results, and a guideline of optimization strategies employed for delivering portable performance on both AMD and NVIDIA GPUs at extreme scale.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 5e0adadd-8341-45e6-9b23-cb9e7b738fc0

Builds on1

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines