SC2023Top-tier venue
5 ExaFlop/s HPL-MxP Benchmark with Linear Scalability on the 40-Million-Core Sunway Supercomputer
Rongfen Lin, Xinhui Yuan, Wei Xue, Wanwang Yin, Jienan Yao, Junda Shi, Qiang Sun, Chaobo Song, Fei Wang
Abstract
HPL-MxP is an emerging high performance benchmark used to measure the mixed-precision computing capability of leading supercomputers. In this work, we present our efforts on the new Sunway that linearly scales the benchmark to over 40 million cores, sustains an overall mixed-precision performance exceeding 5 ExaFlop/s, and achieves over 85% of peak performance, which is the highest efficiency reached among all heterogeneous systems on the HPL-MxP list. The optimizations of our HPL-MxP implementation include the following: (1) a Two-Direction Look-Ahead and Overlap algorithm that enables overlaps of all communications with computation; (2) a multi-level process-mapping and communication scheduling method that uses the entire network as best as possible while maintaining conflict-free algorithm-flow; and (3) a CG-Fusion computing framework that eliminates up to 60% of inter-chip communications and removes the memory access bottleneck while serving both computation and communication simultaneously. This work could also provide useful insights for tuning cutting-edge applications on Sunway supercomputers as well as other heterogeneous supercomputers.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 3a0e1641-0b84-49f1-85fc-c4eed8cb43feRelated papers
- Enabling and scaling the HPCG benchmark on the newest generation Sunway supercomputer with 42 million heterogeneous coresQianchao Zhu, Hao Luo, Chao Yang, Mingshuo Ding et al.SC 2021 · 44 citations
- Climbing the Summit and Pushing the Frontier of Mixed Precision Benchmarks at Extreme ScaleHao Lu, Michael A. Matheson, Vladyslav Oles, J. Austin Ellis et al.SC 2022 · 8 citations
- Scaling the memory wall using mixed-precision - HPG-MxP on an exascale machineAditya Kashi, Nicholson Koukpaizan, Hao Lu, Michael A. Matheson et al.SC 2025 · 3 citations
- Unlocking High Performance with Low-Bit NPUs and CPUs for Highly Optimized HPL-MxP on Cloud Brain IIWeicheng Xue, Kai Yang, Yongxiang Liu, Dengdong Fan et al.SC 2024 · 3 citations
- Acceleration of fusion plasma turbulence simulations using the mixed-precision communication-avoiding krylov methodYasuhiro Idomura, Takuya Ina, Yussuf Ali, Toshiyuki ImamuraSC 2020 · 6 citations
