5 ExaFlop/s HPL-MxP Benchmark with Linear Scalability on the 40-Million-Core Sunway Supercomputer
Rongfen Lin, Xinhui Yuan, Wei Xue, Wanwang Yin, Jienan Yao, Junda Shi, Qiang Sun, Chaobo Song, Fei Wang
摘要
HPL-MxP is an emerging high performance benchmark used to measure the mixed-precision computing capability of leading supercomputers. In this work, we present our efforts on the new Sunway that linearly scales the benchmark to over 40 million cores, sustains an overall mixed-precision performance exceeding 5 ExaFlop/s, and achieves over 85% of peak performance, which is the highest efficiency reached among all heterogeneous systems on the HPL-MxP list. The optimizations of our HPL-MxP implementation include the following: (1) a Two-Direction Look-Ahead and Overlap algorithm that enables overlaps of all communications with computation; (2) a multi-level process-mapping and communication scheduling method that uses the entire network as best as possible while maintaining conflict-free algorithm-flow; and (3) a CG-Fusion computing framework that eliminates up to 60% of inter-chip communications and removes the memory access bottleneck while serving both computation and communication simultaneously. This work could also provide useful insights for tuning cutting-edge applications on Sunway supercomputers as well as other heterogeneous supercomputers.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Enabling and scaling the HPCG benchmark on the newest generation Sunway supercomputer with 42 million heterogeneous coresQianchao Zhu, Hao Luo, Chao Yang, Mingshuo Ding 等SC 2021 · 被引用 44 次
- Climbing the Summit and Pushing the Frontier of Mixed Precision Benchmarks at Extreme ScaleHao Lu, Michael A. Matheson, Vladyslav Oles, J. Austin Ellis 等SC 2022 · 被引用 8 次
- Scaling the memory wall using mixed-precision - HPG-MxP on an exascale machineAditya Kashi, Nicholson Koukpaizan, Hao Lu, Michael A. Matheson 等SC 2025 · 被引用 3 次
- Unlocking High Performance with Low-Bit NPUs and CPUs for Highly Optimized HPL-MxP on Cloud Brain IIWeicheng Xue, Kai Yang, Yongxiang Liu, Dengdong Fan 等SC 2024 · 被引用 3 次
- Acceleration of fusion plasma turbulence simulations using the mixed-precision communication-avoiding krylov methodYasuhiro Idomura, Takuya Ina, Yussuf Ali, Toshiyuki ImamuraSC 2020 · 被引用 6 次
